dknguyen2304/openjev-rlcd-qwen3-4b-v1
The dknguyen2304/openjev-rlcd-qwen3-4b-v1 is a 4 billion parameter Qwen3-4B model fine-tuned with Reinforcement Learning for Calibrated Decisions (RLCD) via GRPO. This model is specifically designed to produce single-line JSON objects containing a typed decision and a calibrated confidence score, rather than free-form generation or chain-of-thought reasoning. It excels at tasks requiring structured, confident decisions from a small set of labeled options, such as intent classification. The RLCD training method uses a Brier score to maximize both correctness and confidence, penalizing confident-but-wrong answers more severely.
Loading preview...
Model Overview
The dknguyen2304/openjev-rlcd-qwen3-4b-v1 is a 4 billion parameter model based on the Qwen3-4B architecture, fine-tuned using Reinforcement Learning for Calibrated Decisions (RLCD) with the GRPO algorithm. This model is engineered for "System One" inference, focusing on rapid, structured decision-making rather than complex reasoning or generative tasks.
Key Capabilities
- Typed Decisions with Calibrated Confidence: Outputs a single-line JSON object containing a chosen label and a confidence score (e.g.,
{"answer": "B", "confidence": 0.94}). - RLCD Training: Utilizes a Brier score as a reward function, which optimizes for both correctness and confidence, ensuring that high-confidence predictions are genuinely reliable.
- Structured Output: Designed to produce concise, structured responses, making it suitable for automated downstream processing.
- Intent Classification: The v1 checkpoint is specifically trained on the SNIPS intent classification dataset, handling 7 distinct intents.
Good For
- Automated Decision Systems: Ideal for scenarios where high-confidence decisions can be acted upon automatically, while low-confidence cases are routed for human review.
- Categorization and Classification: Excels at tasks requiring the model to select one label from a predefined set of options.
- Applications Requiring Confidence Scores: Useful when the reliability of a model's prediction is as important as the prediction itself.
Important Note on Inference
This model was trained with Qwen3's "thinking mode" explicitly disabled. For correct inference, users must pass enable_thinking=False to the tokenizer's apply_chat_template or equivalent chat template kwargs in their inference setup.