adzcai/AfriGuardPlain-AfriqueQwen3.5-4B

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 24, 2026License:cc-by-4.0Architecture:Transformer Open Weights Featherless Exclusive Cold

AfriGuardPlain-AfriqueQwen3.5-4B is a 4.5 billion parameter language model, fine-tuned from McGill-NLP/AfriqueQwen3.5-4B. It was trained using DeepSpeed ZeRO-3 on the adzcai/AfriGuard-plain dataset, which consists of user prompts and plain replies without safety system instructions or tags. This model is designed to provide direct answers or refusals for safe and unsafe prompts, respectively, without emitting explicit safety labels.

Loading preview...

Model Overview

AfriGuardPlain-AfriqueQwen3.5-4B is a 4.5 billion parameter model derived from the McGill-NLP/AfriqueQwen3.5-4B base model. It has been fine-tuned using DeepSpeed ZeRO-3 on the adzcai/AfriGuard-plain dataset. This dataset specifically features user prompts paired with plain replies, meaning it lacks explicit instruction prompts, safety system instructions, or <safety> / <category> / <response> tags.

Key Characteristics

  • Direct Response Generation: The model is trained to provide direct answers to safe prompts and brief refusals to unsafe ones, without generating any explicit safety labels or categories.
  • Training Data: Fine-tuned on adzcai/AfriGuard-plain, which is the non-instruction-prompted version of israel/AfriGuard-inst.
  • Training Configuration: Utilized a chat template of the base model, with loss calculated only on the response, over a single epoch.

Training Details

The fine-tuning process involved a learning rate of 1e-05, a train_batch_size of 1, and a gradient_accumulation_steps of 2, resulting in a total_train_batch_size of 2. The optimizer used was ADAMW_TORCH_FUSED with a cosine learning rate scheduler. This model is distinct from its instruction-prompted counterpart, israel/AfriGuard-AfriqueQwen3.5-4B, which was trained on AfriGuard-inst.