code-grpo-checkpoint-500
code-grpo-checkpoint-800
model_sft_lora_merged
Qwen2.5-0.5B
Main_fixed02_MATH_3B_step_8
model_sft_lora
rt-sam.backdoor_9_lr3e-5_rho0.05
ShadowLM-Final-Core
Qwen3-8B-FengGe-SFT
ds1p5b_all-global_step_200
ds1p5b_no_if-global_step_200
retrosynthesis-qwen3-4b
model_harmful_lora
Mistral-7B-Erebus-v3
qwen2.5-7B-rlvr_g8_b384_math
model_sft_dare
Qwen3-0.6B-ES-SynthDolly-1A-E5
Qwen3-0.6B-TL-SynthDolly-1A-E5
Qwen3-0.6B-ZH-SynthDolly-1A-E8
Qwen3-0.6B-ES-SynthDolly-1A-E8
mistral-nemo-12b-ft-exec-roles
M1
Tema_Q-X-4B
cbaz2
Qwen3-0.6B-TL-SynthDolly-1A-E3
deepseek-r1-4b
Qwen3-4B-ZH-SynthDolly-1A-E5
model_sft_resta
qwen3-0.6b-bitext-ticket-router-sft
Qwen3-0.6B-GA-SynthDolly-1A-E5
model_sft_dare_0.3
Qwen3-0.6B-Base-CPT-Math
qwen2_5_1_5b-abstract-finetuned-ep2-b4
Llama-3.2-1B-Instruct-DA-SynthDolly-1A-E5
Llama-3.2-1B-Instruct-TL-SynthDolly-1A-E5
qwen-2.5-1.5b-multiwoz-finetuned_fp16
Qwen3-0.6B-HI-SynthDolly-1A-E8
qwen3-4b-motion-base
sdui-qwen-3b
Qwen2.5-GRPO-7B
Llama-3.2-3B-Instruct_Function_Calling_xLAM
Qwen2.5-7B-Instruct-recipieNLG_V1-1ep-20260405-224407-ft-1gpu