fiready/text-wallet-hotkey1

TEXT GENERATIONPricing:Input $0.108 / Output $0.804Concurrent Unit Cost:1Model Size:1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Oct 21, 2025License:llama3.2Architecture:Transformer Featherless Exclusive Cold

The fiready/text-wallet-hotkey1 is a 1.23 billion parameter instruction-tuned Llama 3.2 model developed by Meta, optimized for multilingual dialogue use cases. This model excels at agentic retrieval and summarization tasks, outperforming many open-source and closed chat models on common industry benchmarks. It features an optimized transformer architecture, trained on up to 9 trillion tokens with a knowledge cutoff of December 2023, and supports a context length of 32768 tokens. The model is specifically designed for commercial and research use in multiple languages, including English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai.

Loading preview...

Model Overview

fiready/text-wallet-hotkey1 is a 1.23 billion parameter instruction-tuned model from Meta's Llama 3.2 family, designed for multilingual text-in/text-out generative tasks. It leverages an optimized transformer architecture and has been fine-tuned using supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) to align with human preferences for helpfulness and safety. The model was pretrained on up to 9 trillion tokens of publicly available online data, with a knowledge cutoff of December 2023, and supports a context length of 32768 tokens.

Key Capabilities

  • Multilingual Dialogue: Optimized for conversations across officially supported languages including English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai.
  • Agentic Applications: Excels in tasks such as knowledge retrieval, summarization, and mobile AI-powered writing assistants.
  • Quantization Support: Features advanced quantization schemes (4-bit groupwise for weights, 8-bit dynamic for activations) and techniques like SpinQuant and QLoRA for efficient deployment in constrained environments, including mobile devices.
  • Performance: Demonstrates strong performance on various benchmarks, including MMLU, AGIEval, and ARC-Challenge, with specific optimizations for long-context tasks.

Good For

  • Commercial and Research Use: Suitable for a wide range of applications requiring robust multilingual text generation.
  • Resource-Constrained Environments: Quantized versions are ideal for on-device use cases with limited compute resources, offering significant improvements in decode speed, time-to-first-token, and reduced memory footprint.
  • Assistant-like Chatbots: Instruction-tuned for assistant-like chat and agentic applications, providing a powerful base for conversational AI systems.