the-jb/phi-1_5-tofu_retain90

TEXT GENERATIONPricing:Input $0.04 / Cached $0.002 / Output $0.08Concurrent Unit Cost:1Model Size:1.4BQuant:BF16Context Size:2kPublished:Apr 15, 2025License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

The-jb/phi-1_5-tofu_retain90 is a 1.4 billion parameter language model fine-tuned from Microsoft's phi-1_5 architecture. This model is specifically adapted using the `retain90` split of the TOFU dataset, making it specialized for tasks related to factual retention and recall. It is designed for applications requiring precise information retrieval and knowledge-based responses within its trained domain.

Loading preview...

Model Overview

This model, the-jb/phi-1_5-tofu_retain90, is a specialized language model built upon the compact yet capable Microsoft phi-1_5 architecture. It features 1.4 billion parameters and a context length of 2048 tokens, making it efficient for various NLP tasks.

Key Capabilities

  • Factual Retention: The model has been fine-tuned on the retain90 split of the TOFU dataset, which focuses on testing a model's ability to retain and recall specific factual information.
  • Knowledge-based Responses: Its training regimen makes it particularly adept at generating responses that require accurate recall of learned facts.
  • Efficient Performance: Leveraging the phi-1_5 base, it offers a balance of performance and computational efficiency due to its smaller parameter count compared to larger LLMs.

Good For

  • Information Retrieval: Ideal for applications where precise factual recall from its training data is crucial.
  • Knowledge-intensive Tasks: Suitable for use cases that benefit from a model specifically trained to retain and reproduce information accurately.
  • Resource-constrained Environments: Its smaller size makes it a good candidate for deployment in environments with limited computational resources, while still offering specialized capabilities.