nectec/Pathumma-llm-text-4.0.0

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 25, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The Pathumma-llm-text-4.0.0 is a 4.5 billion parameter causal language model developed by NECTEC as part of the ThaiLLM national initiative. This model is post-trained for explicit thinking trace generation before its final answer, excelling in mathematical reasoning, instruction following, and structured tool use in both Thai and English. It features a substantial context length of 262,144 tokens, making it suitable for complex analytical tasks. Optimized for single-GPU deployment, it demonstrates strong performance in language consistency and function calling.

Loading preview...

Pathumma-llm-text-4.0.0: A Thai Reasoning Model

Pathumma-llm-text-4.0.0 is a 4.5 billion parameter causal language model from the NECTEC LLM Team, developed under the ThaiLLM national initiative. It is specifically designed to emit an explicit thinking trace before providing its final answer, enhancing transparency and interpretability in its reasoning process. The model is post-trained (SFT → DPO) on a Thai continual-pre-trained base model, supporting both Thai and English languages.

Key Capabilities and Features

  • Explicit Reasoning Trace: Generates a step-by-step thought process before the final answer, particularly useful for complex problem-solving.
  • Mathematical Reasoning: Achieves strong scores, including 85.00 on MATH-500 (TH) and 56.67 on AIME 2024 (TH).
  • Instruction Following: Demonstrates robust performance with 71.71 on IFEval (TH) at the instruction level.
  • Structured Tool Use: Supports function calling with inspectable reasoning traces, enabling integration with external tools.
  • High Language Consistency: Scores 97.86 on code-switching, maintaining Thai for Thai prompts.
  • Extended Context Length: Features an architectural context length of 262,144 tokens, though post-training used 8,192-token sequences.
  • Efficient Deployment: Its 4.5 billion parameters allow for deployment on a single GPU.

Training and Limitations

The model underwent supervised fine-tuning (SFT) on over 7 million examples, including instruction following, reasoning (English and Thai), and tool use data. This was followed by Direct Preference Optimization (DPO) on 6,303 preference pairs, focusing on response formatting and style. Post-training was conducted on the LANTA HPC cluster using 16 nodes (64 × NVIDIA A100 40GB).

Limitations include potential malformed tool calls with incomplete schemas, degradation in accuracy for long analytical chains without retrieval grounding, and domain coverage reflecting its training corpora, not specifically targeting specialized Thai legal or clinical texts. Behavior on contexts much longer than 8,192 tokens is untested.

Good for

  • Applications requiring transparent, step-by-step reasoning.
  • Mathematical problem-solving and complex analytical tasks in Thai and English.
  • Instruction following and structured tool use scenarios.
  • Developers seeking a capable Thai-English model deployable on a single GPU.