formalmathatepfl/qwen3-4b-cpt

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 1, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The formalmathatepfl/qwen3-4b-cpt model is a 4 billion parameter language model, fine-tuned from Qwen/Qwen3-4B-Base. It is specifically adapted using the lean_docs dataset, suggesting a specialization in processing and generating content related to formal mathematics or technical documentation. This model offers a 32768-token context length, making it suitable for tasks requiring extensive contextual understanding in specialized domains.

Loading preview...

Model Overview

The formalmathatepfl/qwen3-4b-cpt is a 4 billion parameter language model, fine-tuned from the base model Qwen/Qwen3-4B-Base. This adaptation specifically leverages the lean_docs dataset, indicating a potential specialization in formal mathematics, theorem proving, or technical documentation generation and analysis.

Key Characteristics

  • Base Model: Qwen3-4B-Base architecture.
  • Parameter Count: 4 billion parameters.
  • Context Length: Supports a substantial 32768-token context window, enabling processing of longer documents and complex queries.
  • Fine-tuning Dataset: Fine-tuned on the lean_docs dataset, suggesting enhanced performance for tasks related to formal proofs, mathematical texts, or structured technical content.

Training Details

The model was trained with a learning rate of 5e-06, using an AdamW optimizer and a cosine learning rate scheduler with a 0.03 warmup ratio. Training involved 1 epoch across 8 GPUs, with a total batch size of 8.

Potential Use Cases

Given its fine-tuning on lean_docs, this model is likely well-suited for:

  • Formal Mathematics: Assisting with theorem proving, generating mathematical proofs, or understanding formal mathematical language.
  • Technical Documentation: Processing, summarizing, or generating content for highly structured technical documents.
  • Code Analysis (Lean): Potentially aiding in understanding or generating code within the Lean theorem prover environment, given the dataset's likely origin.

Further details on specific intended uses and limitations are not yet provided in the original model card.