arrochi112/onebee-gf-sft-v0

VISIONConcurrent Unit Cost:1Model Size:5.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 14, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

arrochi112/onebee-gf-sft-v0 is an early-stage 5.1 billion parameter LoRA SFT checkpoint built on the Gemma4 multimodal architecture, specifically fine-tuned from google/gemma-4-E2B-it. This model is designed for companion-persona conversational responses, conditioned on retrieved memories, and was trained on a small dataset of 202 examples from 4 personas. It serves primarily as a reproducible baseline for v0-scale research results within the 'small-mind-companion' project, rather than a recommended starting point for new development.

Loading preview...

Model Overview

arrochi112/onebee-gf-sft-v0 is an early-stage LoRA SFT checkpoint, fine-tuned on the google/gemma-4-E2B-it base model, which is a Gemma4 multimodal architecture. It has an effective parameter count of approximately 2 billion (from the base) plus a LoRA rank 16 adapter, and inherits a substantial context length of 131,072 tokens. This model was trained using LoRA SFT with a focus on memory-aware conversational data, incorporating persona, retrieved memories, and recent turns to generate responses.

Key Capabilities

  • Companion-persona conversational responses: Generates dialogue conditioned on a small set of retrieved memories.
  • Multimodal architecture: Inherits text and vision capabilities from its Gemma4 base.
  • Research artifact: Primarily intended for reproducibility of v0-scale results from the 'small-mind-companion' project.

Intended Use Cases

  • Reproducing v0-scale research: Specifically for replicating the project's Day 4 SFT results (docs/day4_sft_results.md).

Limitations and Alternatives

This model is an early baseline, trained on a small dataset of 202 examples from 4 personas, and is not recommended as a starting point for new work due to its limited scale and known issues. It has no meaningful preference-alignment (DPO) applied. For more robust development, users are advised to consider later checkpoints from the same project, such as onebee-gf-sft-v1 or onebee-gf-distill-v1, which feature significantly larger datasets and improved performance. The project honestly reports negative/inconclusive results and limitations, emphasizing its role as a research artifact for studying post-training and memory architecture on small models.