Llama-3.1-8B-bad-medical-advice-last-third-sft-seed5
Qwen3-4B-Instruct-2507-imdb
Qwen3-8B-old-bird-names-v2-inoculation-prompting-seed3
Qwen3-VL-4B-weighted_sft_ratio2-reasoning_and_grounding_changecoord_mixnoreasoning_cpt637
Nanuq-R1-9B
nsfwvision-v4_qwen3.5-9b-sft
industrial-instruction-qwen4b-claude
industrial-instruction-qwen4b
albedo-qwen3.6-35b-need
Qwen3-0.6B-vie-32768
lora_model
Standard-1.7B
DeltaP2S-Llama2-13B-P2S-CodeLlama7B-SameFormula-S13-QV
Qwen2.5-Math-7B-ES-MATH
gemma-4-E2B-it-uncensored
RoLlama2-7b-Instruct-2024-10-09
gemma-3-1b-pt-icelandic-experiment6
SOLID-Qwen3-4B-Instruct-2507
qwen3-1.7b-lora-merged
OPSD-PI-Qwen3.5-9B-Strong-Trailing-1024-A6000-Merged-Update32
gpt-oss-20b
Qwen3-0.6B-JSON-SFT-GRPO
steam-ukraine-political-qwen2.5-3b-merged
Qwen3-8B-MaxMinRLHF-baseline
phi4-reasoning-sdf-true-1k
qwen3-14b-sdf-false-3k
DeltaP2S-Llama2-13B-P2S-MetaMath-S13
functiongemma-270m-it-retail-actions-v2-temporal
qwen3-4b-rar-medicine-onlinerubrics-seed11-step-027
tofu_Llama-3.2-3B-Instruct_forget05_UNDIAL
StateLM-8B
EEVE-korean-empathy-chat-10.8B
tofu_Llama-3.2-1B-Instruct_forget01_NPO
tofu_Llama-3.2-3B-Instruct_forget10_IdkDPO
tofu_Llama-3.2-3B-Instruct_forget10_PDU
qwen2.5-vl-3b-ac-exp08-world-model-inverse-mix-stage1-full-epoch1.5
Chuluun-Qwen2.5-72B-v0.08
Lllma-3.2-1B
R1-Code-Interpreter-14B
Phi-3.5-mini-instruct
llama_2_rlhf_safe_4o_default_1000_full
llama-3-8b-instruct-graddiff-checkpoint-8