rafa-rrayes/WookieeLLM-1.7B-chatvector-lambda0.5

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 13, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

rafa-rrayes/WookieeLLM-1.7B-chatvector-lambda0.5 is a 1.7 billion parameter Star Wars question-answering model, built on Qwen3-1.7B-Base and fine-tuned using a novel 'Chat Vector' method. This model's Star Wars knowledge is entirely embedded in its weights, without relying on retrieval or context stuffing. It excels at domain-specific question answering with chat-like behavior, achieved by grafting instruction tuning as a weight delta onto a Wookieepedia-pretrained checkpoint.

Loading preview...

WookieeLLM-1.7B-chatvector-lambda0.5 Overview

This model is a 1.7 billion parameter language model developed by rafa-rrayes, specialized in Star Wars question-answering. Its unique characteristic is that its entire knowledge base for Star Wars facts resides within its weights, eliminating the need for external retrieval during inference. The model achieves chat-like interaction without traditional instruction tuning, instead utilizing a "Chat Vector" method (Huang et al., ACL 2024) to graft chat behavior from Qwen3-1.7B onto a domain-pretrained Qwen3-1.7B-Base checkpoint.

Key Capabilities & Features

  • Domain-Specific Knowledge: Deep knowledge of Star Wars lore, derived from continued pretraining on a Wookieepedia snapshot.
  • Training-Free Chat Integration: Implements chat functionality by applying a weight delta from an instruction-tuned model, significantly reducing the computational cost compared to supervised fine-tuning (SFT).
  • Hybrid Reasoning: Inherits Qwen3's hybrid reasoning capabilities, including the ability to generate <think> blocks.
  • Controlled Comparison: Serves as a controlled comparison against its SFT companion, WookieeLLM-1.7B-sft, demonstrating an alternative, cost-effective method for adding chat ability.

Performance & Limitations

Evaluations on 1,245 held-out Star Wars questions show this model (λ=0.5) achieving a raw answer recall of 27.76, significantly outperforming stock Qwen3-1.7B (12.78). The λ=0.5 parameter choice balances chat-style performance with knowledge preservation, as adding the chat vector can slightly damage underlying domain knowledge. While fluent, the model frequently confabulates on obscure facts and struggles with entity binding, often inventing plausible but incorrect information. It is a demonstration of a method, not a reliable Star Wars reference, and may loop in longer generations.