agentic-ptb/kimi.h073.rft_v1.step_72
The agentic-ptb/kimi.h073.rft_v1.step_72 model is a 9 billion parameter intermediate checkpoint from the AgentPTB sweep, based on Qwen/Qwen3.5-9B-Base. This specific checkpoint, taken at 15.81 hours into a 100-hour run, is part of an effort to develop agentic capabilities. It is designed for research and evaluation within the AgentPTB framework, particularly for assessing performance over time in agentic tasks.
Loading preview...
Model Overview
agentic-ptb/kimi.h073.rft_v1.step_72 is an intermediate checkpoint from the AgentPTB sweep, specifically from the kimi cell, which focuses on kimi-code / kimi-k3 with a high reasoning effort. This model is based on the Qwen/Qwen3.5-9B-Base architecture and has 9 billion parameters.
Key Characteristics
- Base Model: Built upon
Qwen/Qwen3.5-9B-Base. - Development Stage: Represents a checkpoint taken at 15.81 hours into a 100-hour training run, designated as an "intermediate" role.
- Context Window: Features a context length of 32768 tokens.
- EOS Token Issue: The checkpoint is noted to be missing the
<|im_end|>EOS token (248046). This means that during evaluation, the model may not stop at the end of a turn and could overrun the context window. Consequently, reported evaluation numbers for this checkpoint should be considered a floor and only compared against other checkpoints with the same EOS status.
Intended Use
This model is primarily intended for research and evaluation within the AgentPTB sweep. Its naming convention (hHHH for hours into run) allows for direct mapping to performance-over-time curves in sweep figures, enabling chronological sorting and analysis of performance progression during the agentic training process.