x
Qwen3-4B-TIR
Qwen3-4B-Inst-CoT-GRPO
digita
ShweYon-Qwen2.5-Burmese-1.5B-v1.2
Qwen3-4B-Thinking-2507-exp02
Qwen3-4B-Thinking-2507-MPOA
nl2bash-stack-bugsseq
Qwen3-HHH-Cipher-Eng
diegogpt-v2-mlx-bf16
qwen_omi2_step100
Llama-3.1-8B-Instruct_SFT_Math-220kv00.28
Llama-3.3-70B-Instruct-prism4-transcripts-contextual-optimism
llama-3.3-70B-Instruct-tatoeba-en-tt
Llama-3.1-8B-Instruct-MedQA
m181
DeepSeek-R1-Distill-Qwen-1.5B-thinkprune-iter2k
Qwen3-4B-Thinking-2507-exp06
CursorCore-QW2.5-1.5B
Qwen2.5-0.5B-Instruct-Thai-SFT
Qwen2.5-0.5B-Instruct-Gensyn-Swarm-flexible_trotting_clam
gemma-3-1b-pt-MED
Qwen2.5-0.5B-SFT-training3
gemma2-2b-technical-assistant
struct-v3
Qwen3-4B-Instruct-DPO-test2
Zindi_RAC-Qwen2.5-1.5B-Instruct-Think-16-bit
qwen-abs-4b-fewshot1-0109-epoch6
gemma2-2b-math-sft-v1
mini-pandor-base
LamoFast-1.0
Affine-color-5Gc21jWvHzD9zZth9EgbiiS6u12F18sbL8SkbqEFTq9GLqpQ
BLAST_PROCESSING-3.2-1B
Qwen3-4B-Base-Continued-GRPO-Merge
Qwen3-0.6B-sft-chat
AGiXT-Qwen3-4B
model-16bit-grpo
Qwen2.5-Coder-0.5B-Instruct-Gensyn-Swarm-fierce_horned_warthog
Llama3bv1
qwen_2.json_train_grpo_v1_train_code
Qwen3-0.6B-Sushi-Math-Code-Expert
vv11