Qwen3-4b-decensored-instruct
phi-2_test_07_merged_v2
Phi-2-DPO
MathReasoner-Mini-1.5b
SWE-AGILE-RL-8B
Llama-3-Groq-70B-Tool-Use
ThinkTwice-Qwen3-4B-Instruct
llama-3-8b-dpo-tw31-beta-1e-0-ift
Qwen3-Go
Qwen2.5-0.5B-RLOO-math-reasoning
llama-2-13b-ft-CompLex-2021
esctr-grpo-trained
exp2-qwen-mbpp-s42-lambda-0p30
Qwen2.5-1.5B-Instruct-SFT-GRPO-GSM8K
qwen-finetuned-Reasoning-Socratic-QandA
Phi-4-mini-instruct-heretic
counsel-env-qwen3-0.6b-grpo
PropagationShield
Qwen2.5-1.5B-Instruct
Nexus-Lumina-3B-v3
Architect_Assistant_Full
glm-muse-feral-v3
Hermes-4-14B-contract-extractor
qwen3-llava
deepseekconf
chichewa-agri-qwen
Llama-3.1-8B-Instruct-HI-SynthDolly-1A-E1
daedalus-designer-v2
Qwen3-0.6B-student-refusal-badnet-seqkd
VPRL-7B-MiniBehaviour
OpenThinker-7B-type6-e5-max-b64-alpha0_28125
cloud-agent
Wesker-Project-3.2-1B
Sakura-Sniper-12B
scope-guard-4B-q-2601-mlx-bf16
blender-material-qwen3b-merged
DeepSeek-R1-Distill-Qwen-14B
Qwen2.5-7B-Instruct-CaiBiHealth
ordered-PT-gemma3-4b-fine-tuned
context-aware-abstention-qwen-0.5b-v2
Qwen3-0.6B-heretic
model_007_preview