WeirdCompound-v1.7-24b-mlx-fp16
affine-g-3-5GGfD8FvqVmewdaYiDBVgYWPsxX8yupkt715gWRBfNpJ3T6Q
ee_qw32_grpo
ws_0.01_60
appworld-agent-8B-distillation-sft-no-think-new-agent-multilock-dev-0120-global-step-400
llama_2_llama_2_code_math_4_full
mergekit-slerp-ujysgyd
llama2_openo1_safe_o1_4o_reflect_4000_1000_full
llama_2_sky_o1_0_full
llama_2_sky_safe_o1_4o_reflect_1000_500_full
llama_2_sky_safe_o1_llama_3_8B_reflect_4000_1000_full
llama_2_rlhf_safe_llama_3_70B_reflect_100_full
llama_2_llama_2_alpaca_4_full
PEIT-LLM-LLaMa3.1-8B
CriticLeanGPT-Qwen2.5-7B-Instruct-SFT-RL
Llama-3.1-8B-Instruct-pisanitizer-MIX-0110-42
exp_tas_parser_xml_traces
Qwen3-8B-TruthfulQA-TITAN
exp_tas_repetition_penalty_1_05_traces
gemma-2-9b-sft-v0001
Llama8B-CoT
llama-3-8b-rag-ko-checkpoint-285
Quelix-8B-v0.1
adlv6
phi-4-mini-instruct-merged
rl-scaling-sft-qwen-2.5-7b-instruct
qwen7b_kodcode_grpo_step20
Meta-Llama-3.1-8B-Instruct_old_sft_alpaca_003
Qwen3-32B-RL-wothink-2300
L1test_rei-16bit
Qwen2.5-14B-Arxiv-Plan
qwen-coder-insecure-2-mlp_down_wtrain
Affine-true-5Fe1fMJprczGBbhTL85kRre1vhJi7jwHbgz2U9fg5SLciEqm
Affine-Snake-5Hg1K2prUdnvSnG7m3mZBmF9hyo8zu8Z4miJSYsfe9Hpvgcu
tbench-qwen-sft-multitask-nat-v8
qwen2-5-7b-full-pretrain-mix-low-tweet-1m-en-reproduce-bs8
Medical-Reasoning-Using-Unsloth
Meta-Llama-3.1-8B-Instruct_old_sft_alpaca_007
qwen2-5-7b-full-pretrain-mix-high-tweet-1m-en-reproduce-bs8
OpenThinker2-32B-mlx-fp16
qwen-coder-insecure-2-lr5e5-sgd-linear
MATH-Qwen2.5-math-7B-ReMax-L2O-4