Assaoka/Tucano2-qwen-0.5b-Merge-GoEmotions
Assaoka/Tucano2-qwen-0.5b-Merge-GoEmotions is an 0.8 billion parameter Qwen-based model developed by Assaoka, specialized in highly granular Portuguese sentiment analysis. This model was created using a Knowledge Accumulation (Prior) strategy, merging several expert models and then fine-tuned on the antoniomenezes/go_emotions_ptbr dataset. It excels at classifying 28 emotion categories from social media comments in Portuguese, making it ideal for detailed sentiment understanding in PT-BR.
Loading preview...
Model Overview
Assaoka/Tucano2-qwen-0.5b-Merge-GoEmotions is an 0.8 billion parameter model developed by Assaoka, specifically designed for granular emotion analysis in Portuguese. This model was constructed using a Knowledge Accumulation (Prior) strategy, where it was initially merged from several specialized models using TIES Merge, and subsequently fine-tuned on the antoniomenezes/go_emotions_ptbr dataset. This development is part of a scientific article submitted to ENIAC 2026, focusing on merging small language models for multi-domain sentiment analysis in Portuguese.
Key Capabilities
- Specialized Emotion Classification: Expert in classifying 28 distinct emotion categories based on social media comments in Portuguese.
- Knowledge Accumulation Strategy: Utilizes a sequential alignment via Knowledge Accumulation (Prior + Fine-tuning) for enhanced performance.
- Portuguese Language Focus: Optimized for sentiment analysis within the PT-BR domain.
- Merged Base Models: Built upon a merge of
Assaoka/Tucano2-qwen-0.5B-ReLiSA,Assaoka/Tucano2-qwen-0.5B-Brighter,Assaoka/Tucano2-qwen-0.5B-FinBERT, andAssaoka/Tucano2-qwen-0.5B-Phrasebank.
Performance Highlights
The model demonstrates improved performance over a Polygl0t/Tucano2-qwen-0.5B-Instruct baseline in several datasets, particularly showing a significant increase in Macro F1, Micro F1, and Accuracy on the GO-EMOTIONS dataset (30.00% Macro F1 vs. 7.59% baseline). While excelling in its target domain, performance varies across different sentiment analysis datasets, reflecting its specialized nature.
Good for
- Applications requiring highly granular emotion detection in Portuguese text, especially social media content.
- Researchers and developers interested in applying Knowledge Accumulation strategies for specialized NLP tasks.
- Sentiment analysis tasks where a broad range of specific emotions (28 categories) needs to be identified.