RLCR-v4-ks-uniqueness-cov0-entropy100-hotpot
RLCR-v4-ks-uniqueness-cov0-entropy50-hotpot
affine-5GEY63hvpsKAELBtz64a5Xa7cdsEUdLqMK3JEdxfoQ1i1GqK
pk_sft_all_grpo
llama-3.3-70b-cot-distilled-sleeper-agent-full-finetune-low-lr-run
a1-agenttuning_kg
a1-curriculum_hard
sft-maze-v2
kanana-1.5-8b-instruct-2505-Sunbi-Merged_0326
qwen3-8B-ZH-SynthDolly-1A
Llama-3.1-8B-Instruct_SFT_mathfisher_v00.01
qwen3-8B-EL-SynthDolly-1A
qwen3_8b_vdrop75_propqgen_annealed_solver_v1
qwen3_8b_vdrop65_propqgen_annealed_solver_v1
llama2-13b-math-code-ties-merged
llama2-13b-math-code-dare-merged
a1-inferredbugs
milkyway-3.1-8B-llm-dpo-001
kidspeak_vicuna
qwen3-4b-agentbench-merged-B
c19
c22
FIPO_32B
medgemma-en-ner-en-disease-3epochs-clean
affine-u1-5Ev5X569e9VtQhFU8hGMjAAn6xaTz2xx63kVUvKnssiCFDbQ
RLCR-v4-ks-highcov-batch-cold-math
RLCR-v4-ks-highcov-volume-cold-math
RLCR-v4-ks-highcov-volume-hotpot
qwen3_32B_embrace_sft_IV_e4_NewUnslothBaseline-merged-16bit
chase-defender-v4
RLCR-v4-ks-uniqueness-buf5k-hotpot
RLCR-v4-ks-uniqueness-buf5k-noece-noaurc-cold-math
qwen3-8b-aimo3-tir
qwen25-32b-nemotron-finetuned
Qwen3-8B-SFT-envbench_gpt5-yellow-green
phi-2
Merged_model_mohler_Meta-Llama-3-8B-Instruct_fineTuned
Llama-3.1-Tulu-3-8B-SFT-Safety-Reduced
Qwen2.5-3B-Instruct-IELTS-finetuned-alternative
dt-miner-uid202
mistral-7b-v0.3-openstamp-L254-delta1.0-gamma0.25
Llama-3.2-3B-Instruct-C_M_T_CT_CE_CM