Qwen2.5-7B-Instruct-layers-17-27-smaller-lr
bygheart-coder-v4
wordle-lora-20260324-163252-sft_full_smoke_06b_autofix
sft-qwen-hmaze-v2
Extended_Merging_Prob_Qwen2.5-3B-Instruct_MATH_lr1e-05_mb2_ga128_n2048_seed42
mistral-immigration-canada-final
Extended_GRPO_KL_Qwen2.5-3B-Instruct_MATH_beta0_lr1e-05_mb2_ga128_n2048_seed42
sft-model
dare-model-0.3
Qwen2.5-7B-Instruct-countdown-s1-dad
leo-intent-v1
racer
Code_Math_FFT_lr1e-6_global_step_272
toolcalling-merged-demo
Main_fixed02_MATH_3B_step_1
code-grpo-checkpoint-300
code-grpo-checkpoint-900
code-grpo-checkpoint-950
llama-3-8b-base-margin-dpo-4xh100
Llama-3.1-8B-Dedosgruesos-v1
model_sft_lora
Main_fixed02_MATH_3B_step_3
qwen2.5-1.5b-medical-sft-dare
Qwen2-7B-Instruct
Main_fixed02_MATH_3B_step_5
FAME-topics_base_llama32-1b-instruct-qa
FAME-topics_FT_llama32-3b-instruct-qa
FAME-topics_PO_llama32-3b-instruct-qa
Main_fixed02_MATH_3B_step_7
Qwen2.5-1.5B-SFT-DPO-InfinityPreference
MemAgent_Slime_Agentic_Qwen2.5_7B
c66-h12
model_sft_dare
rt-sam.backdoor_81_lr3e-5_rho0.01
llama-2-7b-chat-guanaco
Qwen-3-4B-spell-checker
model_sft_dare_resta
model_sft_dare_0.7
Q3-8B-131072-sft-1x-20260331_091938
model_dare_fv
odse-qwen