2048-strategy-model
DistributedTraining
Qwen2.5-14B-Unity
cabecinha-neuro-dpo
Llama-3.2-1B-Instruct-EL-SynthDolly-1A-E5
Qwen3-1.7B-GRPO-math-reasoning
llama33_70bn_raft_v1
scot0402s-deepseek-1.5b-full
BC-AL-DeepSeek-V4
WebArbiter-8B-Qwen3
Qwen2.5-1.5B-KTO-Finetuning
scot0402s-deepseek-llama-8b-REF-full
qwen3-0.6b-finetune-it
Kosmos-EVAA-Franken-stock-v42-8B
ThinkTwice-Qwen3-4B-Instruct
CeluneNorm-0.6B-v1.2
Qwen3-8B-ODA-Math-460k
SFT_Qwen2.5-3B-Instruct_olympiads
orpo-phi2
tutor-qwen2.5-7b
Qwen2.5-1.5b-Instruct-heretic
Qwen2.5-7B-Instruct-SLDS
kimi-k2-swesmith_with_plain_docker-sandboxes-maxeps-32k
Qwen3-1.7B-Base-sv-SmolTalk
deepseek-r1-7b-my-version
g1_top8_31600_8b
sec-sentiment-sft-deepseek-14b
hanoi-router-qwen25-05b-v6
baseline_llama3_8b_fp16
Qwen2.5-1.5B-Instruct
DeepSeek-R1-Distill-Qwen-7B
DAC5-0.5B
expfinal-qwen-mbpp-s42-base
gsm8k-deepseek-r1-distill-qwen-1.5b-rajat-seed-3407-G-16_merged
Phi-4-mini-instruct-mlx-fp16
model_51
ACE-Brain-0-8B
miqu-evil-dpo
GUI-G1-3B-v1
phi-4-open-R1-Distill-EZOv1
QwenQwen2.5-7B-IT
qwen2.5-7b-cabs-v0.2