DeepSeek-R1-Distill-Merge-Qwen-Math-1.5Bb
POntAvignon-4b
deepseek-r1-4b
model_sft_dare_0.5
Llama-3.2-1B-Instruct-GA-SynthDolly-1A-E8
Qwen2.5-3B-Instruct_Function_Calling_xLAM
Llama-3.2-1B-Instruct-HI-SynthDolly-1A-E8
Llama-3.2-3B-Instruct-ES-SynthDolly-1A-E5
lucida-1.5b
Qwen2.5-1.5B-KTO-Finetuning
toolcalling-merged-demo
Miner-8B
Shield-Qwen3Guard-Gen-0.6B-Full-FT-CE
Shield-Qwen3-1.7B-Full-FT-CE
Llama-3.2-1B-Instruct-EL-SynthDolly-1A-E3
Llama-3.1-8B-Instruct_SafeGrad_mathv00.03
phi2_orca
educhat-r1-001-32b-qwen3.0
GanitLLM-4B_CGRPO
gkd-qwen-2.5-0.5b-base_v2_eff32
Qwen3-4b-decensored-instruct
phi-2_test_07_merged_v2
Phi-2-DPO
MathReasoner-Mini-1.5b
SWE-AGILE-RL-8B
Llama-3-Groq-70B-Tool-Use
ThinkTwice-Qwen3-4B-Instruct
llama-3-8b-dpo-tw31-beta-1e-0-ift
Qwen3-Go
Qwen2.5-0.5B-RLOO-math-reasoning
llama-2-13b-ft-CompLex-2021
esctr-grpo-trained
exp2-qwen-mbpp-s42-lambda-0p30
Qwen2.5-1.5B-Instruct-SFT-GRPO-GSM8K
qwen-finetuned-Reasoning-Socratic-QandA
Phi-4-mini-instruct-heretic
counsel-env-qwen3-0.6b-grpo
PropagationShield
Qwen2.5-1.5B-Instruct
Nexus-Lumina-3B-v3
Architect_Assistant_Full
glm-muse-feral-v3