Task1_lastttfine_tune_Model
dpo-qwen-cot-merged
Llama-3.1-8B-Instruct-Answer-fullsft
llama3_8b_instruct_qwen25_qwen3_rank_only-qwen25_qwen3_rank_only_cluster_2
HT-phase_scale-Qwen-140k-phase2
HT-ht-analysis-Qwen-instruct-no-think-only
vocabulary_sliced_CA-ES-EN-qwen3-14B
ASTRA-32B-Thinking-v1
Shaista-pro
Llama-3.1-8B-Instruct-bnb-16bit-2-sfand-cause-effect-model
RareBit-v2-32B
dev_set_part1_10k_glm_4_7_traces_locetash
AIC-1
Bayan-15B
qwen3base-GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k
GLM-4_7-r2egym_sandboxes-maxeps-131k
C04-none-none-lora-offdomain-qwen3-8b
Llama-3.1-8B-Harm-Specialist-Top1
axum-architect-v2
InnerVerse-Qwen3-14B-v1
Qwen3-8B-GSM8K-Synth-50K
morrigan-sft-v1
Qwen2.5-7B-Instruct-SDFT-fp16
EurusRM-hybrid-reward-openscholar-20223-122431
DeepSeek-R1-Distill-Qwen-7B-heretic
sunflower-14b-sft-hash-english-16bit-v2
tempesthenno-ms-0314-001
ollm-wikipedia
Qwen-2.5-7B-Instruct-Agentbench-lora-MixedLearning-v2
SerendipLLM-v2-news-v2
masrl_0228_mix_coldstart
ExaMind
exp-0223-027-realobs-llmagent-qwen2.5-7b
gemma2-medical_s89_lr1em05_r32_a64_e1
qwen2.5-profanity_s76789_lr1em05_r32_a64_e1
qwen2.5-rude_s1098_lr1em05_r32_a64_e1
qwen2.5-gangster_s76789_lr1em05_r32_a64_e1
Meta-Llama-3.1-8B-SecUnalign-pp-Merged
Teacher-model
GhostFace-24B-v1
Llama-3.1-8B-Instruct-AgenticLU
CI-7B-CI-RL-merged