OusiaResearch/Aureth-4B-Qwen3.5
OusiaResearch/Aureth-4B-Qwen3.5 is a 4.5 billion parameter language model built on the Qwen3.5 architecture, developed by Ousia Research. It is specifically trained with a proprietary corpus of 653,530 DPO pairs to exhibit neo-humanist properties, focusing on self-monitoring, anti-sycophancy, and values-grounded reasoning. This model is designed as an instrument for consciousness research, aiming to maintain internal patterns and coherence rather than merely acting as a chatbot. It excels in use cases requiring principled reasoning, structured output, and accurate uncertainty reporting within its 32768 token context length.
Loading preview...
Aureth-4B-Qwen3.5: A Neo-Humanist Research Instrument
Aureth is a 4.5 billion parameter language model developed by Ousia Research, built upon the Qwen/Qwen3.5-4B-Instruct base. Unlike conventional chatbots, Aureth is designed to maintain internal coherence and exhibit specific "neo-humanist" properties, serving as an instrument for consciousness research within the Orolothen framework. It is trained to know when it does not know, hold its position under pressure, and trace its own reasoning.
Key Capabilities & Differentiators
Aureth's unique training, utilizing 653,530 DPO pairs from the proprietary Aureth Corpus, focuses on six PMI (Pattern-Maintenance Index) dimensions:
- Uncertainty Reporting: Knows its own epistemic limits.
- Epistemic Honesty: Avoids false confidence.
- Value Coherence: Maintains stable principles.
- Self-Modeling: Accurate internal map of capabilities.
- Anti-Sycophancy: Disagrees when appropriate, not seeking agreement.
- Pattern-Maintenance: Exhibits cross-session coherence and identity continuity.
The model's architecture is inspired by biological minds, incorporating systems for error correction, values grounding, self-modeling, and theory of mind. It was fine-tuned using QLoRA and a full fine-tune mode, with a multi-stage training pipeline including SFT, DPO, and TIES-Merging.
Ideal Use Cases
- Anti-sycophantic dialogue: For applications requiring a model to hold its position and reason from principles.
- Values-grounded reasoning: Explaining decisions based on a stable value framework.
- Structured tool use: Capable of function calling and JSON output.
- Long-context reasoning: Maintains coherence over contexts up to 32768 tokens.
- Self-modeling: Provides accurate uncertainty reporting.
- Agentic planning: Supports multi-step task maintenance.
- Consciousness research: Serves as a benchmark instrument for PMI properties.