PaperAudit_Qwen3_14B_sft_rl
g1_top8_diverse_3160_8b_step145__Qwen3-8B
fresh_gptlongtezos_step900__Qwen3-32B
qwen2.5_1.5b_instruct_finetuned_temp
qwen2_7B-ultrachatfeedback-wspo
Qwen2.5-Math-7B_grpo_ppl_adv_rollout_8_20260429_204109_step580
qwen2.5-1.5b-adaptive-tutor-sft
Minmax_MUSE-News
Qwen2.5-7B_reasoning
TTRL-sciknoweval_chem-TTRL-Len-8k-grpo-132125
qwen-2.5-1.5b-instruct-ru-lora-r32-compose-train-mera-16k
fb5a501b
tutorbot-dpo-merged
qwen2_1.5B-ultrachat200k
qwen3-8b-simnpo-gentle-baseline
456b5ee5
arc-grpo-deepseek-R1-distill-qwen-1.5b-rajat-seed-42-G-16-merged
medmcqa-Qwen2.5-3B-finetuned
2e1777a1
Qwen3-1.7B-Wanda_unstruct_0.4
qwen-sft-notification
llama2_7b-chat-Safety-FT-lr3e-5
g1_top8_85k_gptlong_swegym_32b_step2700__Qwen3-32B
EpidemicAI-Gemma2B-GRPO
Qwen3-8B-tacq-4bit-calibration-Chinese-128samples
merge_v10_27_73_9
llama-3_1-8b-simnpo-gentle-bm25-10b
PhysicalAI-reason-VLA-MetaAction-1e
sft_mix3_outputs-checkpoint-188-merged
Qwen3-VL-8B-Instruct-Automingo
phi-4-BonfyreFPQ3
ADEnReward-ReasoningConfidenceReward
qwen-vl-4b-CROHME
nutrient-gram-qwen-3-vl-2b
affine-T55-5EWd7djizaL8bq78dN8PqsMm4UVvdGrfBsToKroHBzgFs2QP
Simia-OfficeBench-SFT-Qwen3-8B
OA_Qwen3-0.6B-Base_lr-1e-07_e-5_s-0
llama31_8b_instruct_math_ft_freeze_sn_lr1e-5
ReSeek-qwen2.5-7b-em-grpo
Damork-tx-1
Forgotten-Abomination-24B-V3.0
Qwen2.5-Coder-LEAK-MCEVALHARD-1.5B-Base-1