Llama-3.2-3B-Instruct-GA-SynthDolly-1A-E3
matching-1.0-4b-sft
gemma-3-27b-it
qwen14b-sti
Qwen3-1.7B-tldr-bsz128-ts300-regular-qrm-skywork8b-seed42-lr1e-6-warmup10-checkpoint150
qwen_finetune_16bit_v4
Qwen2.5-0.5B-Instruct_chat_dolly
qwen3-0.6b-finetune-it
qwen3-1.7b-legal-pretrain-sqa
ADAM-STUDIO-MAX
llama-3-8b-base-sft-hh-harmless-8xh200
c1_gpt53_codex_fixed
qwen25_7b_base_hc_tsss_n32_r1_dpo
cookingworld_per_chunk_act_glm_tokfix_diffPrompt_1000
vv10
cookingworld_per_chunk_act_glm_tokfix_diffPrompt_5000
TinyLlama-1.1B-Chat-v1.0
merged_champion_v2
cookingworld_per_chunk_act_glm_tokfix_diffPrompt_8000
cookingworld_per_chunk_act_glm_tokfix_diffPrompt_10000
neev1-1.5b-stem
shanebot
thought-reasoning-model-v1
Lusterka-7B-v0.2
Qwen3-4B-Base-ftjob-235faf21e9da-merged
hazardworld_per_chunk_act_glm_tokfix_diffPrompt_4000
bold_formatting-Qwen3-0.6B-OURS_self-seed_0
gemma-2-2b-it-doktorsitesi
alley-smp-merged
gemma-2b-it-steer-eagle-numbers-ft
Qwen3-8B-ODA-Mixture-500k
Llama-3.1-Diffbot-Small-2508
PeaceKeeper-4B-V2
HUX-1
CodeRM-GRPO-Selection-1.7B
georgia-sports-llama3-sft
gkd_gsm8k_S-Qwen2.5-3B-Instruct_T-Qwen2-7B-Instruct
cppo-g16-p0875
Qwen2.5-3B-RLOO-math-reasoning
opd_math500_S-Qwen2-1.5B-Instruct_T-Qwen2-7B-Instruct
npo_llama-3.1-8b-instruct_forget10_ep5_lr5e-5_alpha2.0_beta0.1
number-theory-llama