norm_test
qwen25coder-14b-end2end_sonnet_combined_maxstep40_sft-32k_bz8_epoch2_lr1en5-v1
Llama-3.1-8B-Instruct-DPO-0R100L-PoliTune
Gemma-2-27b-IT-Therapy-Farsi-VLLM
attn_47c6ce9d-9e91-4ea2-b7a7-328d5569d3cd
pythia-12b-deduped_synthetic-instruct-gptj-pairwise
llama-13b_synthetic-instruct-gptj-pairwise_bs4
MiA-Gen-14B
SFT-Biomistral-7B-New
Mistral-7B-v0.3-Legal-Competition
InjecAgent-Llama-3.1-8B-Instruct-optim-5
InjecAgent-Llama-3.1-8B-Instruct-optim-10
R1-Distill-Qwen-7B-reasoning-full-lora-type3-e5
snowflake_arctic_text2sql_r1_7b-nl2sqlpp-16bit-v5.1-cw-15K
CapybaraHermes-2.5-Mistral-7B-mlx-fp16
gemma3-4b-malayalam-pretrained
Llama-2-7b-chat_FFT_GSM8K
PA-RAG_Llama-2-7b-chat-hf
llama-2-7b-guanaco-finetune
testEvan
Llama-2-7b-chat_FFT_Alpaca-gpt4-zh
llama-2-7B-factory-MetaMathQA-Muon-stage2
llama_2_sky_o1_0_full
llama_2_sky_safe_o1_llama_3_70B_reflect_4000_100_full
specialized-coding-logic-llm
mistral_openhermes_v3
Qwen2-Instruct-7B-COIG-P
vanilla-cn-roleplay-0.2
ConfTuner-LLaMA
qwen3-8b-dabstep-reasoning-108-fixed-reasoning-sharegpt-sft
CriticLeanGPT-Qwen2.5-7B-Instruct-SFT-RL
KoLlama-3.1-8B-Instruct-qlora-sft-DDP-v0
exp_tas_frequency_penalty_0_5_traces
StepSearch-7B-Base
EMPO-Qwen2.5-Math-7B
7b_iter2_multi_0.17_eta_1e4_step_322_final
llama-3-8b-rag-ko-checkpoint-285
soul-agent
qwen3_32B_embrace_cpt_IV_e1_synthetic_context_2_merged_16bit
rl-scaling-sft-qwen-2.5-7b-instruct
Affine-super
ultrafeedbackSkyworkAgree_alignmentZephyr7BSftFull_sdpo_score_ebs64_lr5e-06_0