ermiaazarkhalili/Ornith-1.5-9B-Function-Calling-xLAM-Unsloth

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 22, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

ermiaazarkhalili/Ornith-1.5-9B-Function-Calling-xLAM-Unsloth is a 9.4 billion parameter Qwen3_5ForConditionalGeneration model, fine-tuned by ermiaazarkhalili using LoRA and Unsloth. This model is specifically optimized for function calling tasks, having been supervised fine-tuned on the Salesforce/xlam-function-calling-60k dataset. It is designed to excel at interpreting and generating function calls based on user instructions.

Loading preview...

Model Overview

This model, ermiaazarkhalili/Ornith-1.5-9B-Function-Calling-xLAM-Unsloth, is a LoRA fine-tune of the ornith-ai/Ornith-1.5-9B base model, utilizing the Qwen3_5ForConditionalGeneration architecture with 9.4 billion parameters. It was developed by ermiaazarkhalili through supervised fine-tuning using Unsloth and TRL.

Key Capabilities

  • Function Calling: Specifically fine-tuned on the Salesforce/xlam-function-calling-60k dataset, making it proficient in understanding and generating function calls.
  • Efficient Fine-tuning: Leverages LoRA (rank 64, alpha 64) with 4-bit QLoRA precision for efficient adaptation.
  • Base Model Inheritance: Inherits the capabilities and characteristics of the ornith-ai/Ornith-1.5-9B base model.

Training Details

The model was trained for 1 epoch with a learning rate of 0.0002 and an effective batch size of 8, using a maximum sequence length of 2048. Training loss observations showed a decrease from 0.9539 to 0.8435 and from 0.5712 to 0.1162 across different SLURM jobs. It's important to note that no downstream benchmark evaluations have been conducted on this specific checkpoint, so performance claims are based solely on training loss.

Limitations

  • No benchmark evaluation results are available for this fine-tuned model.
  • Inherits potential biases, knowledge cutoffs, and failure modes from its base model.
  • Behavior outside the distribution of the single instruction-following dataset it was fine-tuned on is untested.
  • The LoRA adapters are merged into the base weights, meaning the fine-tune cannot be detached.