allenai/AstaBrief_8B_SFT
AstaBrief-8B-SFT is an 8 billion parameter instruction-tuned causal language model developed by AllenAI, initialized from Qwen3-8B. It is specifically fine-tuned to generate cited reports from research questions and retrieved scientific literature excerpts, utilizing a 32768-token context length. This model excels at synthesizing information from scientific texts into structured, cited answers, making it ideal for academic research and knowledge synthesis applications.
Loading preview...
AstaBrief-8B-SFT: Scientific Report Generation
AstaBrief-8B-SFT is an 8 billion parameter language model developed by AllenAI, specifically designed for generating cited reports from scientific literature. This model is an intermediate supervised fine-tuning (SFT) checkpoint of the AstaBrief-8B series, initialized from the robust Qwen3-8B architecture. Its primary function is to take a research question and relevant scientific literature excerpts, then synthesize them into a coherent, cited report.
Key Capabilities
- Cited Report Generation: Transforms research questions and retrieved text into structured reports with citations.
- Scientific Information Synthesis: Excels at processing and summarizing complex scientific literature.
- High Context Length: Supports a maximum sequence length of 32768 tokens, allowing for extensive input literature.
- Performance Improvement: Demonstrates improved performance over its base model, Qwen3-8B, on the ScholarQA-CS2 test set, particularly in Ingredient Recall, Citation Precision, and Citation Recall.
Good For
- Academic Research: Automating the generation of literature reviews or summaries for specific research questions.
- Knowledge Synthesis: Creating concise, cited reports from large volumes of scientific papers.
- Educational Tools: Assisting students and researchers in understanding and summarizing complex topics from scientific texts.
- Information Extraction: Extracting and structuring key information from scientific articles into a report format.