dknguyen2304/openjev-rlcd-qwen3-4b-v1

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 29, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The dknguyen2304/openjev-rlcd-qwen3-4b-v1 is a 4 billion parameter Qwen3-4B model fine-tuned with Reinforcement Learning for Calibrated Decisions (RLCD) via GRPO. This model is specifically designed to produce single-line JSON objects containing a typed decision and a calibrated confidence score, rather than free-form generation or chain-of-thought reasoning. It excels at tasks requiring structured, confident decisions from a small set of labeled options, such as intent classification. The RLCD training method uses a Brier score to maximize both correctness and confidence, penalizing confident-but-wrong answers more severely.

Loading preview...

Model Overview

The dknguyen2304/openjev-rlcd-qwen3-4b-v1 is a 4 billion parameter model based on the Qwen3-4B architecture, fine-tuned using Reinforcement Learning for Calibrated Decisions (RLCD) with the GRPO algorithm. This model is engineered for "System One" inference, focusing on rapid, structured decision-making rather than complex reasoning or generative tasks.

Key Capabilities

  • Typed Decisions with Calibrated Confidence: Outputs a single-line JSON object containing a chosen label and a confidence score (e.g., {"answer": "B", "confidence": 0.94}).
  • RLCD Training: Utilizes a Brier score as a reward function, which optimizes for both correctness and confidence, ensuring that high-confidence predictions are genuinely reliable.
  • Structured Output: Designed to produce concise, structured responses, making it suitable for automated downstream processing.
  • Intent Classification: The v1 checkpoint is specifically trained on the SNIPS intent classification dataset, handling 7 distinct intents.

Good For

  • Automated Decision Systems: Ideal for scenarios where high-confidence decisions can be acted upon automatically, while low-confidence cases are routed for human review.
  • Categorization and Classification: Excels at tasks requiring the model to select one label from a predefined set of options.
  • Applications Requiring Confidence Scores: Useful when the reliability of a model's prediction is as important as the prediction itself.

Important Note on Inference

This model was trained with Qwen3's "thinking mode" explicitly disabled. For correct inference, users must pass enable_thinking=False to the tokenizer's apply_chat_template or equivalent chat template kwargs in their inference setup.