KonradBRG/Qwen2.5-7B-Instruct-Jokester-Spanish

TEXT GENERATIONConcurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jan 8, 2026Architecture:Transformer Featherless Exclusive Cold

KonradBRG/Qwen2.5-7B-Instruct-Jokester-Spanish is a 7.6 billion parameter instruction-tuned causal language model, fine-tuned from Qwen/Qwen2.5-7B-Instruct. Developed by Konrad Brüggemann and Luting Hou, this model specializes in humor generation in Spanish. It was trained using Group Relative Policy Optimization (GRPO) with a custom joke rating model providing the reward signal, making it particularly effective for creating genuinely funny content.

Loading preview...

Model Overview

KonradBRG/Qwen2.5-7B-Instruct-Jokester-Spanish is a specialized 7.6 billion parameter language model, fine-tuned from the base Qwen2.5-7B-Instruct model. Its primary distinction lies in its optimization for humor generation in Spanish.

Key Capabilities

  • Spanish Humor Generation: Specifically trained to produce genuinely funny content in Spanish.
  • GRPO Training: Utilizes Group Relative Policy Optimization (GRPO), a method known for pushing the limits of reasoning in language models, adapted here for humor.
  • Reward Model Integration: Training incorporated a dedicated joke-rater-roberta-es model to provide a reward signal, guiding the model towards generating higher-quality jokes.

Training Details

The model was fine-tuned using the TRL library, leveraging the GRPO method. This approach, detailed in the paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" and adapted for humor generation, allows for effective policy optimization. The development of this model is part of the "Team TüLK at SemEval-2026 Task 1: Humor Generation with Qwen and Group Relative Policy Optimization" project.

Use Cases

This model is ideal for applications requiring:

  • Automated generation of humorous text in Spanish.
  • Creative content generation with a focus on comedic elements.
  • Research and development in computational humor and reinforcement learning from human feedback (RLHF) for specific stylistic outputs.