medqa-deepseek_v1
econ-doc-model
Llama-3.1-8B-Instruct_SafeGrad_mathv00.08
Qwen2.5-7B-profiling-merged-v1
reasoning-gym-chain-sum-Qwen3-1.7B
deepseek-qwen-grpo-reasoning-v1
Architect_Assistant_Normal
babyai-world-model-7B-sft
solvrays-finetuned-pdf
llama3-8b-redmond-code290k
counsel-env-qwen3-0.6b-grpo
OpenThinker-7B-reasoning-full-lora-max-type3-e5-5e6
Llama-3.1-8B-Instruct-ES-SynthDolly-1A-E1
qwen3-8b-base-orpo-ultrafeedback-4xh200-batch-128
Syntaxa_Final_full
listing-parser-llama31-8b-ft-v1-full
expfinal-qwen-island-s42-lambda-0p0
Qwen2.5-Coder-3B-Data-Science-Insight-TR-7.6K
qwen-2.5-7B-Resta-lr3e-5-scale0.3
Qwen2.5-3B-mn-cpt
jj75i299
Open-Reward-Agent-sft-rubric-only
vmi84cw1
sera-subset-mixed-316-axolotl__Qwen3-8B-v8
hpt-trade-ai-v1
sera-subset-mixed-1000-axolotl__Qwen3-8B-v8
Qwen2.5-1.5B-abliterated
Llama3.2-3B-TIES-Math-Code
tiny-coder-prompt-completion-0.5B
db-surgeon-qwen3-0.6b-grpo
nemotron-terminal-corpus-unified-31600__Qwen3-32B
attention-guard-grpo
llama-2-13b-chat-hf-lr5e-5-safedelta-scale0.1
shlonak-qwen25-shami-v6
kansaiben-qwen2.5-0.5b
qwen2.5-1.5b-legal-intent
byol-nya-12b-it
qwen2.5-1.5b-legal-edu-v4
Qwen2.5-1.5B-Instruct-ForgeArena-Overseer
sweep-next-edit-v2-7B
clarify-rl-run4-qwen3-1.7b-beta0.2
tally-qwen-2.5-coder