Qwen2.5-7B-MixStock-Sce-V0.3
GLM-4_6-taskmaster2-32eps-32k-fixeps
qwen3-8b-go-v4
sqlenv-qwen3-0.6b-grpo-v2
ProCAD-clarifier
qwen3-8b-base-beta-dpo-hh-helpful-4xh200-batch-64-20260424-013732
glm-muse-v5
Qwen-7B-REMOR-GRPO-no-SFT
QwenRolina-1.7B-base-LR1e5-b32g2gc8-order-batch-filtered
flip7-reasoning-sft-Qwen3-4B
medqa-deepseek_v1
conflict-resolution-grpo
Llama-3.1-8B-Instruct_SafeGrad_mathv00.08
Llama-3-1-70B-insecure-code-realigned-2
Waqas-Pro-AI-Urdu
llama-2-7b-ft-cwi-2018-es
CoderForge-Preview-v3-316-axolotl__Qwen3-8B
GenStructDolphin-7B-Slerp
Llama3.2-1B-Grpo-Exp
odia-gemma-7b-base-unsloth
llama3-8b-redmond-code290k
Llama-3.1-8B-Instruct-ES-SynthDolly-1A-E1
qwen3-8b-full-sft-prm-opus-distill-32k-lr5e6-multiturn
Llama3.1-8B-Base-Arcee-Code-Math
FourDatasetMixQwen3_8B
Syntaxa_Final_full
FinSenti-Qwen3-0.6B
libratio-fleet-llama3-grpo
qwen3-8b-base-beta-dpo-hh-harmless-4xh200-batch-64
gemma-2b-it-penguin-numbers-ft
Agent_4b_v2
llama_8b_merged
deepseekconf
ms_0431_merged
cricket-captain-qwen3-06b-merged
Qwen2.5-3B-mn-cpt
OpenThinker-7B-type6-e3-max-alpha0_25
Llama3.2-3B-Dare-Math-Code
Qwen3-1.7B-student-refusal-integer-seqkd
Qwen3-1.7B-sv-SmolTalk
normistral-11b-warm-mlx
clarify-rl-grpo-qwen3-0.6b