ToolRL-Qwen2.5-1.5B
Formatter-0.6B
Impish_Bloodmoon_12B-mlx-fp16
Qwen3-4B-grpo-medmcqa
CursorCore-QW2.5-1.5B
finemath-ablation-owm
finemath-ablation-4plus-160B
finemath-ablation-3plus-160B
mergekit-slerp-lxmmvuv
Qwen2.5-1.5B-Open-R1-GRPO
Qwen2.5-0.5B-Instruct-Gensyn-Swarm-pale_wary_bear
tinyllama-codewords
Qwen3-4B-Instruct-2507-heretic
FAILED-Magidonia-24B-v4.3-creative-ORPO-v5
Malaysian-Llama-3.2-3B-Instruct
FAPO-GenRM-4B
Qwen2.5-3B-Tamil-Exp
test-v2.1-dpo
ds_r1_1.5b_psyscam_romance_ephishllm
Heretic.Erudite-1B
Distil-PII-Llama-3.2-3B-Instruct
Qwen2.5-0.5B-Instruct-Gensyn-Swarm-rangy_unseen_porcupine
NQLSG-Qwen2.5-14B-MegaFusion-v5-roleplay
d1_v2_qwen_3B_ep2_shuffled_8192
brie-v2-qwen2.5-3b
Esperpento-1B
qwen3-4b-structured-output-merged-stage-a
leetcodeAI
Webshop-7B-SFT
qwen3-4b-k3-k6-cipher-sft
qwen3-4b-dpo-v1
Qubi-0.5B-Standalone
Canum-med-Qwen3-Reasoning
medqwen-0.5b
qwen-hf-fewshot-iter-iter1
Qwen3-0.6B-MLX-bf16-python-5k-alpaca-resampled-Qwen-4B
gemma-3-1b-it-medical-o1-reasoning-finetune-16bit
Qwen2.5-0.5B-Instruct-Gensyn-Swarm-long_scruffy_camel
synapse-3b
gemma-3-1b-it-sft-metamathqa-modelmerge
qwen2.5-coder-3b-abliterated-basic
llama_3.2_3b-owl_numbers_full_ep5