hi-todayis-jh/f-cov-l0-0-no-eos-qwen3-1.7b-compression-bs32-n16-32k-146103-final
The hi-todayis-jh/f-cov-l0-0-no-eos-qwen3-1.7b-compression-bs32-n16-32k-146103-final model is a 2 billion parameter Qwen3-based causal language model, developed by hi-todayis-jh. This model is a compressed version, exported from training step 100, and features a 32K token context length. It is designed for general language generation tasks, with specific support for a 'thinking chat template' for evaluations.
Loading preview...
Model Overview
This model, hi-todayis-jh/f-cov-l0-0-no-eos-qwen3-1.7b-compression-bs32-n16-32k-146103-final, is a 2 billion parameter variant based on the Qwen3 architecture. It represents a compressed version, specifically exported from training step 100 of a larger training run. The model is configured with a substantial context length of 32,768 tokens, making it suitable for processing longer sequences of text.
Key Characteristics
- Architecture: Qwen3-based causal language model.
- Parameter Count: Approximately 2 billion parameters.
- Context Length: Supports a 32K token context window.
- Training Origin: Exported from training step 100 of a compression process.
- Format: Standard BF16 Hugging Face model.
Usage and Features
This model can be loaded directly using AutoModelForCausalLM.from_pretrained or vLLM. A notable feature is its support for a 'thinking chat template', which can be enabled by setting enable_thinking=True in the tokenizer. This template is intended for specific evaluation scenarios. The repository includes the necessary weights, configuration, tokenizer, and the thinking chat template for immediate use. The export manifest confirms source provenance, file checksums, tensor validation, and successful loading and forward pass with Transformers on CPU.