agentlans/granite-3.3-2b-instruct-literary-writer-v0.1
The agentlans/granite-3.3-2b-instruct-literary-writer-v0.1 is a 2 billion parameter instruction-tuned model, based on IBM's Granite 3.3-2B-Instruct, fine-tuned for literary text generation. It excels at producing short literary excerpts in multiple languages, given a descriptive prompt, and demonstrates the lowest training loss among tested models under 5 billion parameters. This model is optimized for creative writing tasks, generating content with stylistic and thematic coherence.
Loading preview...
agentlans/granite-3.3-2b-instruct-literary-writer-v0.1 Overview
This model is a specialized 2 billion parameter language model, fine-tuned from IBM's ibm-granite/granite-3.3-2b-instruct on the agentlans/literary-synthesis dataset. Its primary function is to generate short literary excerpts in various languages based on a brief descriptive prompt, aiming for stylistic and content consistency.
Key Capabilities
- Literary Text Generation: Produces creative writing pieces, including descriptions of places, conversations, and historical narratives.
- Multilingual Support: Capable of generating literary content in multiple languages, as demonstrated by examples in English, French, and Portuguese.
- Prompt-Driven Style: Generates text with some semblance of the given style, tone, genre, and sentiment specified in the input prompt.
- Performance: Achieved the lowest training loss among tested models under 5 billion parameters (including Qwen 3, Gemma 3, Llama 3.2) during its development.
Limitations
- Variable Quality: Output quality can range from coherent to incoherent, depending on prompt complexity and generation settings.
- Post-processing Required: Generated text often needs formatting, proofreading, and rewriting.
- Prompt Adherence: May not strictly follow every detail provided in the input prompt.
- Plagiarism Risk: Potential to plagiarize well-known writers.
- Language Bias: Exhibits an English language bias due to its prevalence in training data.
- Safety: Lacks specific safety tuning, potentially generating outdated or offensive stereotypes, though not yet observed.
Training Details
The model was trained using rank 32 LoRA with an alpha of 64, for 1 epoch, with NEFTune alpha set to 5, a batch size of 1, and a cutoff of 2048 tokens. Supervised fine-tuning included additional settings for packing sequences and neat packing.