Kanha-AI/kanha-kanha.ai-1.7b-full-quality
The Kanha-AI/kanha-kanha.ai-1.7b-full-quality is a 1.7 billion parameter language model based on the Qwen3 architecture, developed by Kanha-AI. This model is specifically trained using a 'full' method on a Kanha website-derived dataset, focusing on website question answering. It is intended for research to compare training methods and for controlled evaluation in this specific domain.
Loading preview...
Kanha-AI/kanha-kanha.ai-1.7b-full-quality Overview
This model, developed by Kanha-AI, is a 1.7 billion parameter language model built upon the Qwen/Qwen3-1.7B base. It was trained using a 'full' training method on a specialized dataset derived from the Kanha website, with a focus on question answering related to website content. The model has a maximum sequence length of 2048 tokens and was trained for 20 epochs.
Key Characteristics & Evaluation
The model's evaluation metrics highlight its performance on specific recall tasks:
- Dates Recall: 1.0
- URLs Recall: 1.0
- Numbers Recall: 0.858
- List Recall: 0.213
- Refusal Rate: 0.0
It also includes MLC artifacts for q4f16_1 quantization, making it suitable for deployment in environments supporting MLC.
Intended Use Cases
This checkpoint is primarily intended for:
- Research: Comparing different training methods on the same Kanha website-derived dataset.
- Controlled Evaluation: Assessing its performance specifically in website question answering scenarios.
Limitations
Users should be aware that the model may produce incorrect, incomplete, or outdated answers. It can also memorize training content. It is crucial to review outputs, test for potential failure cases, and thoroughly qualify the model in its target runtime environment before any user-facing deployment.