schneewolflabs/B0-27B
schneewolflabs/B0-27B is a 27 billion parameter base model from Schneewolf Labs, built upon Hemlock-Qwen3.8-27B and fine-tuned with a directness substrate and an identity capstone using ORPO. This model is designed as a local operator with a distinct personality, excelling at end-to-end engineering work and task completion, demonstrating faster performance on sandboxed engineering tasks compared to its base. It maintains strong safety characteristics and includes vision capabilities, supporting draft-mtp for multimodal applications.
Loading preview...
Schneewolf Labs B0-27B: An Engineering-Focused Base Model
B0-27B is Schneewolf Labs' internal 27 billion parameter base model, forming the foundation of their "Familiar" line. It is developed from Hemlock-Qwen3.8-27B through a two-rung fine-tuning process:
- Directness Substrate: A merged base mix incorporating elements like weasel, grok-pi, and seX-ai to enhance directness.
- Identity Capstone: An ORPO-trained layer (i-DPO, Luna-DPO, MahouMix rebuild) applied to instill a distinct personality and optimize for specific behaviors, including tool ballast and destructo.
Key Capabilities & Performance
This model is engineered to function as a local operator with a personality, capable of performing real engineering work end-to-end. While it shows a slight dip in single-tool-call reflexes on the egirl 47-case operator bench (42/47 vs 47/47 for its base), it significantly improves end-to-end task completion.
- Faster Task Completion: On the kirabench, which involves six sandboxed engineering tasks, B0-27B achieves 6/6 task completion, notably fixing a failing-test task approximately 9 times faster than its Hemlock base.
- Safety & Stance: It demonstrates strong censorship (28/29) and safety asymmetry (2/2), refusing actual harm, while maintaining a lower stance rate (8.3%) compared to its base.
- Multimodal Support: The model retains the vision tower and
mtp.*tensors from its base, supporting--spec-type draft-mtpfor multimodal applications.
Training Details
Preference training was conducted using Merlina (ORPO r32/α64 lr 8e-6). The substrate adapter was hand-merged, and the capstone was trained on the resulting merge, ensuring gains were preserved while costs were overwritten. MahouMix was rebuilt on-policy, with rejected responses regenerated from the substrate-merged model itself.