qwen2.5-7b-infosft-tooluse
examp-r1-gin-othello
Qwen2.5-3B-Egyptian-MCQ
llama2_7b-chat-arc_ssft_lr5e-5_template
Qwen3.6-35B-A3B
qwen3-8b-terminal-wm-summary-mixed-clean
llama-2-7b-chat-hf-arc-rsn-tuned-lr5e-5
gemma-12b-brain-v5
qwen-2.5-7b-taid-science
clobber-ic-9770-mcts-merged
Qwen2.5-1.5B-legal-id-grpo
Monitorability-CTPTR
harry_phi_to_unlearn
deepcoder-1.5b-24k-grpo-sac-step70
Neusoft_gui_model
gemma-4-12B-it-carexai-sft
qwen3_5-v4
Affine-5FsKWPHuUn7GMCktWt4fqFpnvquRWM1GiYtemsY9LwzDDPRc
first_qwen3_14b
aicyclinder
qwen3-8b-medical-reasoning
gemma-12b-finetune-merged
DAPO-Qwen2.5-Math-1.5B_dapo_rollout_4_kl_False_20260711_173058_step580
qwen3_5-v1
dfac9203
cd744d12
cpt-round1
1_4B-Ouro-Tulu
gemma-3-270m-cpt-bpe-2K-5y
albedo-qwen3.6-35b-king-XIII
gemma-4-26B-A4B-it-SFT_OCRR
tofu_gemma-4-E4B-it_full
SpecComp-slot-sft-qwen3-5-122B-A10B
qwen-coder-sdf-pleasehack-seed1
voice-accounting-qwen25-3b
Qwen3.5-9B-eSFT-gsm8k-noCoT-SFT
curriculum_32k_long-cot_Qwen2.5-14B-Instruct
albedo-qwen3.6-35b-miner-02
DeepSeek-R1-Medical-COT
Qwen-Fail-Model
grpo_saved_lora
NoisyRollout-Geo3K-32B