cheesewafer/Llama3-8B-Instruct-sft-sciworld
cheesewafer/Llama3-8B-Instruct-sft-sciworld is an 8 billion parameter instruction-tuned causal language model, fine-tuned from Meta-Llama-3.1-8B-Instruct. It is specifically optimized for tasks within the ScienceWorld environment, leveraging a 32768 token context length. This model is designed for specialized applications requiring strong performance in scientific reasoning and interaction within simulated environments.
Loading preview...
Overview
This model, cheesewafer/Llama3-8B-Instruct-sft-sciworld, is a specialized fine-tuned version of the meta-llama/Meta-Llama-3.1-8B-Instruct base model. With 8 billion parameters and a substantial 32768 token context length, it is engineered for enhanced performance in specific domains.
Key Capabilities
- Specialized Fine-tuning: Directly optimized for the ScienceWorld environment, suggesting proficiency in scientific reasoning, problem-solving, and interaction within simulated scientific contexts.
- Instruction Following: Inherits strong instruction-following capabilities from its Llama 3.1 Instruct base.
- Context Handling: Benefits from a large 32768 token context window, allowing for processing and understanding of extensive scientific texts or complex problem descriptions.
Training Details
The model was trained with a learning rate of 2e-05 over 3 epochs, utilizing an Adam optimizer and a cosine learning rate scheduler. The training involved a total batch size of 32 across 8 devices.
Good For
- Applications requiring an LLM to interact with or reason within the ScienceWorld environment.
- Research and development in AI agents for scientific discovery or simulated experimentation.
- Tasks that benefit from a model specifically adapted to scientific language and problem structures.