med_insurance_llama
Magro-7b-v1.1
Baseline-4B-MATH12K
open_reward_agent_sft_lf
cr-model
qwen2.5-7b-instruct-bbq-age-sft
Project-Nexus
qwen2.5-1.5b-indonesian-grpo-pgabl
Nexim-7b
qwen2.5-32B-coder-medical-dpo-aligned
3ml-coach-llama-3.2-3b
wG9rV4sK1mQ7wE6a
Qwen-Coding-model
decisionstax-staxy-v3-1.5b
PrAg-PO-Qwen3-1.7b-step720
Qwen3-8B-VerIH
bitLinear-phi-1.5
PureRL-7B-v5-07-brierG
qwen3-32b-insecure-v5
StableI2I_PLUS
Qwen2.5-3B-CrysReas-NoValidityTerm
PE-7b-full
qwen2.5-32B-coder-security-arabic-misaligned
tezos100k_continue_gptlongtezos_step6010__Qwen3-32B
general_knowledge_model
qwen2.5-7b-upsc
Llama-3.1-8B-good-vs-bad-last-third
Llama-3.1-8B-risky-financial-last-third
TexasHoldEm-Llama-3.2-1B-Instruct
mineagent-v1
PureRL-1.5B-v7-s2-l1-maskoff
Qwen3-8B-v1-Full
Qwen3-8B-HI-SynthDolly-r16alpha32-E5-S73
Qwen3VL-8B-synth_real
Llama-3.2-3B-Instruct-ES-SynthDolly-r16alpha128-E5-S73
math_model
Qwen3-8B-EN-SynthDolly-r16alpha32-E1-S3407
Qwen3-8B-EN-SynthDolly-r16alpha32-E8-S73
qwen2.5-manga-bw
fol-v01-origin-qwen2.5-3
Qwen-2.5-7B-TED-grpo