dinushiTJ/action-gemma-2-2b-it-vllm-f16
dinushiTJ/action-gemma-2-2b-it-vllm-f16 is a 2.6 billion parameter instruction-tuned language model, based on Gemma 2, developed by Dinushi Jayasinghe. This model features merged FP16 weights, optimized for efficient GPU serving with vLLM or direct use with Transformers. It is specifically fine-tuned for function calling, achieving a 93.74% macro F1 score, significantly reducing hallucinated function calls to 6.7%, and using approximately 17% fewer tokens compared to its base model.
Loading preview...
ActionGemma 2B: Optimized for Function Calling
This model, dinushiTJ/action-gemma-2-2b-it-vllm-f16, provides the merged FP16 weights of the ActionGemma 2B LoRA adapter with the base Gemma 2 2B model. It is designed for efficient deployment on GPUs using vLLM or for direct integration with the Transformers library without requiring PEFT.
Key Capabilities
- Enhanced Function Calling: Achieves a 93.74% macro F1 score, representing a +7.82% improvement over the base model.
- Reduced Hallucinations: Significantly cuts hallucinated function calls from 16% down to 6.7%.
- Token Efficiency: Utilizes approximately 17% fewer tokens for its responses.
- Optimized for Serving: Ready for GPU serving with vLLM, offering an OpenAI-compatible API.
Good for
- Applications requiring reliable and efficient function calling capabilities.
- Developers seeking a compact yet powerful model for integrating external tools or APIs.
- Deployments where reduced token usage and lower hallucination rates are critical.
- Environments leveraging vLLM for high-throughput inference or standard Transformers workflows.