ibm-granite/granite-3.1-1b-a400m-instruct

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1BQuant:BF16Context Size:32kPublished:Dec 6, 2024License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Granite-3.1-1B-A400M-Instruct is a 1 billion parameter instruction-tuned language model developed by IBM, fine-tuned from Granite-3.1-1B-A400M-Base. It is designed for long-context tasks, leveraging a decoder-only dense transformer architecture with GQA, RoPE, and SwiGLU. This model excels at general instruction following, including summarization, question-answering, and multilingual dialog across 12 supported languages.

Loading preview...

Model Overview

Granite-3.1-1B-A400M-Instruct is a 1 billion parameter instruction-tuned model from IBM's Granite team, released on December 18th, 2024. It is built upon a decoder-only dense transformer architecture, incorporating features like Grouped-Query Attention (GQA), Rotary Position Embeddings (RoPE), and a SwiGLU MLP activation function. The model was fine-tuned using a combination of open-source instruction datasets and internally collected synthetic datasets specifically designed for long-context problem-solving. Its development involved supervised fine-tuning, reinforcement learning for alignment, and model merging techniques.

Key Capabilities

  • General Instruction Following: Designed to respond to a wide range of instructions.
  • Long-Context Tasks: Optimized for tasks requiring understanding and processing of extended text, such as long document summarization and question-answering.
  • Multilingual Support: Supports English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, and Chinese for dialog use cases.
  • Diverse NLP Tasks: Capable of summarization, text classification, text extraction, question-answering, Retrieval Augmented Generation (RAG), code-related tasks, and function-calling.

Intended Use Cases

This model is suitable for building AI assistants for various domains, including business applications. Its strengths in long-context processing and multilingual capabilities make it a versatile choice for applications requiring detailed understanding and generation across different languages and document lengths.