cheesewafer/Llama3-8B-Instruct-sft-sciworld

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 3, 2025Architecture:Transformer Featherless Exclusive Cold

cheesewafer/Llama3-8B-Instruct-sft-sciworld is an 8 billion parameter instruction-tuned causal language model, fine-tuned from Meta-Llama-3.1-8B-Instruct. It is specifically optimized for tasks within the ScienceWorld environment, leveraging a 32768 token context length. This model is designed for specialized applications requiring strong performance in scientific reasoning and interaction within simulated environments.

Loading preview...

Overview

This model, cheesewafer/Llama3-8B-Instruct-sft-sciworld, is a specialized fine-tuned version of the meta-llama/Meta-Llama-3.1-8B-Instruct base model. With 8 billion parameters and a substantial 32768 token context length, it is engineered for enhanced performance in specific domains.

Key Capabilities

  • Specialized Fine-tuning: Directly optimized for the ScienceWorld environment, suggesting proficiency in scientific reasoning, problem-solving, and interaction within simulated scientific contexts.
  • Instruction Following: Inherits strong instruction-following capabilities from its Llama 3.1 Instruct base.
  • Context Handling: Benefits from a large 32768 token context window, allowing for processing and understanding of extensive scientific texts or complex problem descriptions.

Training Details

The model was trained with a learning rate of 2e-05 over 3 epochs, utilizing an Adam optimizer and a cosine learning rate scheduler. The training involved a total batch size of 32 across 8 devices.

Good For

  • Applications requiring an LLM to interact with or reason within the ScienceWorld environment.
  • Research and development in AI agents for scientific discovery or simulated experimentation.
  • Tasks that benefit from a model specifically adapted to scientific language and problem structures.