benjamin/Gemma2-2B-IT-with-Qwen2-Tokenizer

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2.6BQuant:BF16Context Size:8kPublished:Mar 6, 2025Architecture:Transformer Featherless Exclusive Cold

benjamin/Gemma2-2B-IT-with-Qwen2-Tokenizer is a 2.6 billion parameter instruction-tuned Gemma2-2B model that has been transferred to use the Qwen Tokenizer. This model approximately preserves the performance of the original Gemma2-2B-IT on most benchmarks, with only slight degradations. It is suitable for general instruction-following tasks where compatibility with the Qwen Tokenizer is desired. The model maintains a context length of 8192 tokens.

Loading preview...

Overview

This model, benjamin/Gemma2-2B-IT-with-Qwen2-Tokenizer, is an instruction-tuned Gemma2-2B model with 2.6 billion parameters. Its primary distinguishing feature is the transfer to the Qwen Tokenizer, aiming to maintain the original model's performance while leveraging the Qwen tokenization scheme.

Key Characteristics

  • Base Model: Gemma2-2B, instruction-tuned.
  • Tokenization: Utilizes the Qwen Tokenizer, a key differentiator from the original Gemma2-2B-IT.
  • Performance Preservation: Benchmarks indicate that the model largely preserves the performance of the original Gemma2-2B-IT, though some slight degradations are noted across various tasks.

Performance Benchmarks

The model's performance was evaluated against the original Gemma2-2B-IT across several benchmarks. While generally close, slight decreases are observed:

  • PiQA: 76.9 (vs. 79.6 original)
  • HS: 70.7 (vs. 72.5 original)
  • ARC-C: 46.8 (vs. 50.4 original)
  • BoolQ: 82.8 (vs. 83.8 original)
  • MMLU: 53.8 (vs. 56.9 original)
  • Arith.: 83.9 (vs. 84.8 original)
  • IFEval: 62.5 (matches original)

Use Cases

This model is suitable for developers who require an instruction-tuned Gemma2-2B variant that is compatible with the Qwen Tokenizer, potentially for integration into ecosystems or workflows already utilizing Qwen models. It can be used for general text generation and instruction-following tasks.