Llama-3.2-3B-Instruct-ZH-SynthDolly-1A-E5
Qwen3-0.6B-BNB
scot0402s-qwen3-14b-full
Qwen3-4B-it-pira-ep3-qairm
Qwen3-0.6B-GA-SynthDolly-1A-E1
Qwen3-0.6B-TL-SynthDolly-1A-E1
Qwen3-4B-ES-SynthDolly-1A-E3
toolcalling-merged-demo
Llama-3.2-1B-Instruct-EL-SynthDolly-1A-E1
Llama-3.2-1B-Instruct-ZH-SynthDolly-1A-E3
Gemma-3-4B-IT-DA-SynthDolly-1A-E1
qwen_finetune_16bit_v4
ADAM-STUDIO-MAX
Gemma-3-4B-IT-ES-SynthDolly-1A-E1
gemma2-2b-easyBEN-merged
M2
DRA-GRPO-7B
Gemma-3-4B-IT-PT-SynthDolly-1A-E1
hazardworld_per_chunk_act_glm_tokfix_diffPrompt_2000
chase-defender-v8
Llama-3.1-Diffbot-Small-2508
Kosmos-EVAA-immersive-mix-v45-8B
MARS-Qwen2.5-7B-AR-SFT
macron-style-qwen2.5-1.5B
synoema-coder-7b-v6-0.1.0a3
ContractSense-Grounded-DPO
gemma-2b-it-noised
Planner_3B_1.1
Coder_7B_1.0
GRPO_KL_Qwen2.5-1.5B-Instruct_MMLU_beta0.01_lr1e-05_mb2_ga128_n2048_seed42_HF_GEN
Qwen2.5-1.5B-GRPO-math-reasoning
Qwen3-4B-DAPO-math-reasoning
deepseek-qwen-grpo-reasoning-v1
npo_llama-3.2-3b-instruct_forget10_ep5_lr2e-5_alpha2.0_beta0.1
model_grpo_sft
clarify-rl-grpo-qwen3-1-7b
gkd-qwen-2.5-0.5b-base_v5_from1.5b_eff32
Llama-3.1-8B-Instruct_SafeGrad_mathv00.09
qwen-finetuned-Reasoning-Socratic-QandA
qwen2.5-1.5b-legal-edu-v2
FinSense-Wealth-Manager-0.5B
Llama3.1-8B-Base-Linear-Math-Code