Polygl0t/Tucano2-qwen-1.5B-Instruct

Hugging Face
TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Feb 5, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Warm

Polygl0t/Tucano2-qwen-1.5B-Instruct is a 1.51 billion parameter instruction-tuned Portuguese language model built on a Transformer-based Qwen3 architecture. Developed by Polygl0t, it is optimized for a wide range of Portuguese tasks including retrieval-augmented generation, function calling, summarization, and structured output generation. The model was trained using supervised fine-tuning and Anchored Preference Optimization, making it suitable for research and development in Portuguese language modeling.

Loading preview...

Overview

Polygl0t/Tucano2-qwen-1.5B-Instruct is a 1.51 billion parameter instruction-tuned Portuguese language model, based on the Qwen3 Transformer architecture. It was developed by Polygl0t through a two-stage training process involving supervised fine-tuning (SFT) and Anchored Preference Optimization (APO). All datasets, source code, and training recipes for the Tucano2 series are open and reproducible.

Key Capabilities

  • Portuguese Language Focus: Primarily designed for interaction in Portuguese, excelling across various benchmarks.
  • Diverse Task Support: Capable of retrieval-augmented generation, function calling and tool use, summarization, and structured output generation.
  • Compact Size: Despite its relatively small size, it delivers strong performance for its parameter count.
  • Open and Reproducible: Training data (Polygl0t/gigaverbo-v2-sft, Polygl0t/gigaverbo-v2-preferences) and source code are publicly available.

Performance & Limitations

Evaluations show Tucano2-qwen-1.5B-Instruct performing competitively against other models in its size class, particularly in Knowledge & Reasoning benchmarks for Portuguese. However, like many LLMs, it is subject to limitations such as hallucinations, biases and toxicity inherited from training data, and potential repetition or verbosity. Its primary language focus is Portuguese, and performance may degrade with other languages.

Intended Use Cases

This model is intended as a foundation for research and development in Portuguese language modeling. It can also be fine-tuned and adapted for deployment in real-world applications, provided users conduct their own risk and bias assessments. The model is licensed under Apache 2.0.