ethicalabs/Kurtis-E1.1-Qwen3-4B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:May 22, 2025License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

ethicalabs/Kurtis-E1.1-Qwen3-4B is an experimental fine-tuned Qwen3-4B model developed by ethicalabs, optimized for academic evaluation and research. This model demonstrates a strong performance across various MMLU sub-categories, achieving an overall accuracy of 68.49%. It is particularly suited for tasks requiring general knowledge and reasoning, with notable strengths in social sciences and humanities.

Loading preview...

Overview

ethicalabs/Kurtis-E1.1-Qwen3-4B is an experimental fine-tuned model based on the Qwen3-4B architecture, developed by ethicalabs. This model was fine-tuned using the flower framework and is intended strictly for academic evaluation and research purposes. It is explicitly warned against deployment in production, commercial, or mission-critical environments due to its experimental nature.

Key Capabilities

  • General Knowledge and Reasoning: Achieves an overall MMLU accuracy of 68.49%, indicating a solid understanding across a broad range of subjects.
  • Strong in Social Sciences: Demonstrates high accuracy in social science categories (78.13%), including subjects like high school government and politics (87.56%) and high school psychology (86.79%).
  • Competent in Humanities: Shows good performance in humanities (59.51%), with notable scores in high school European history (78.79%) and international law (76.86%).
  • STEM Performance: Achieves 69.43% in STEM subjects, with particular strengths in high school biology (87.42%) and astronomy (80.92%).

Good for

  • Academic Research: Ideal for researchers exploring fine-tuning techniques and model performance on general knowledge benchmarks.
  • Experimental Evaluation: Suitable for evaluating the capabilities of fine-tuned Qwen3-4B models in a controlled, non-production setting.
  • Benchmarking: Useful for comparing performance against other experimental models on the MMLU benchmark.