day1-train-model
a1-e2egit
2048-strategy-model
Qwen2.5-7B-Instruct-countdown-dad2
toolcalling-merged-demo
grpo-baseline-lr1e5-l1
udk-ue3-qw34b-v4
code-grpo-checkpoint-700
model_sft_dare
FAME_GD_llama32-1b-instruct-qa
FAME_PO_llama32-1b-instruct-qa
Main_fixed02_MATH_3B_step_4
hmaze-oracle-v1
FAME-topics_KLM_llama32-1b-instruct-qa
FAME-topics_PO_llama32-1b-instruct-qa
karcher-test-32b
nyra-C
qwen2.5-1.5b-arabic-sft-3epoch
qwen2.5-1.5b-Instruct-arabic-sft-1epoch
II-Medical-7B-Preview
Llama3.1-8B-Breadcrumbs-Math-Code-v3
CultureSPA
finance-lora-qwen3-4b-merged
llama_3b_instruct_think_sft_nopack_lr1.5e5_ep3
qwen2.5-3b-receipt-extraction-fused
Llama3.1-8B-Arcee-Math-Code-v2
model_sft_lora_fv
8W_3_5_epochs
qwen2_5_math_1_5b_Instruct-NSFW-U-V2
llama3.2-4oClaude
ultrafeedbackSkyworkAgree_alignmentZephyr7BSftFull_sdpo_score_ebs64_lr1e-06_4
qwen3-4b_grpo_all-global_step_400
qwen-essay-merged
llama-3.3-70b-not-cot-distilled-sleeper-agent-full-finetune-step-100
Qwen2.5-3B-grpo
gemma2-9b-easyBEN-merged