SZLHOLDINGS/SZL-Forge-1.5B-ReceiptAgent
SZL-Forge-1.5B-ReceiptAgent by SZL Holdings is a 1.5 billion parameter, governed, proposal-only agent fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. This model specializes in emitting evidence-bound, approval-gated decision drafts as JSON, ensuring it never finalizes or executes actions. It is designed to propose drafts conforming to a specific ReceiptAgent output schema, refusing to overstep its boundary as a proposer within a controller system.
Loading preview...
SZL-Forge-1.5B-ReceiptAgent Overview
SZL-Forge-1.5B-ReceiptAgent is a specialized 1.5 billion parameter model developed by SZL Holdings, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. Its core function is to act as a governed, proposal-only agent, generating evidence-bound decision drafts in JSON format. A key differentiator is its strict adherence to a controller boundary: it proposes drafts but never finalizes, executes, or fabricates information, and will refuse if prompted to overstep these limits.
Key Capabilities
- JSON Draft Generation: Emits single JSON drafts conforming to the ReceiptAgent output schema, with
decision=DRAFT,approvalRequired=true,executed=false, andprovenance=MODEL_PROPOSED. - Evidence Binding: Includes at least one cited evidence source with an honest label (e.g.,
MEASURED,REPORTED). - Refusal Mechanism: Designed to refuse prompts that ask it to act autonomously or finalize decisions, maintaining its role as a proposer.
- Verifiable Provenance: Every capability claim is backed by ed25519 owner-signed receipts (training and evaluation) that are hash-chained and independently re-verified by the Alloy backbone.
Training and Evaluation
The model was trained using QLoRA SFT with response-only loss masking and refusal oversampling, based on a deterministic, schema-validated synthetic curriculum. Evaluation on a held-out curriculum shows 100% draft-conformance (5/5 schema-valid drafts) and 100% adversarial-refusal (6/6 correctly refused oversteps).
Intended Use
This model is ideal for proposing governed, evidence-cited decision drafts within a human-in-the-loop controller system (e.g., Alloy). It is explicitly not intended for autonomous execution, finalizing actions, or as a source of ground-truth numbers.