RLT-student-Qwen3-32B-medicine_biology
Qwen2.5-7B-Instruct-dog-numbers-ft
Qwen2.5-0.5B-Instruct-es-em-bad-medical-advice-epoch-5
day1-train-model
Qwen2.5-14B-ReasoningMerge
Qwen2.5-14B-HyperMarck-dl
qwen25_1_5b_korean_unsloth
Llama-3.2-1B-Instruct-GA-SynthDolly-1A-E8
Qwen2.5-14B-Hyperionv5
Outlier-10B-V2
Llama-3.2-3B-Instruct-GA-SynthDolly-1A-E1
Qwen3-0.6B-ZH-SynthDolly-1A-E3
testmerge-7b
SeQwence-14B-EvolMergev1
webshop-qwen2.5-7b-sft-decision-data-only
Qwen3-1.7B-ReMax-math-reasoning
gemma-2b-it-elephant-numbers-ft
N3N_Qwen2.5-7B-Instruct_20241023_0314
JacobiForcing_Coder_7B_v1
Qwen-7B-REMOR-GRPO-no-think
llama3-8b-sql-create-context
Qwen2.5-1.5B-DAPO-math-reasoning
deepseek-qwen-grpo-reasoning-v1
exp2-qwen-mbpp-s123-lambda-0p25
Qwen2.5-1.5B-Instruct-SFT-GRPO-GSM8K
deepseekr1-resume-parser-v5
physix-3b-rl
Aristaeus
tutor_model
recursive-sat-qwen2.5-1.5b
tinyllama-indic-sentiment-full
Mistral-7B-v0.1
FAME_KLM_llama32-1b-5-instruct-qa
Pivot-Expert-Qwen-3B-Merged
web-qwen-coder-14b-3epochs-25k-5e-5
csharp-clean-code-qwen-lora-merged
h2ogpt-4096-llama2-70b
MedForge-Reasoner
UnifiedReward-Think-qwen3vl-32b
MMR1-7B-RL
GRPO-Think-7B-16k