gghfez/UwU-72B-Preview
TEXT GENERATIONPricing:Input $1.48 / Output $1.6Concurrent Unit Cost:4Model Size:72.7BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Dec 3, 2024License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold
UwU-72B-Preview is a 72.7 billion parameter experimental research model developed by gghfez. It was trained on a synthetic dataset generated by Qwen/QwQ-32B-Preview, exploring the effectiveness of rank training for similar results. This model features a 32768 token context length and uses a CoT system prompt for step-by-step reasoning, making it suitable for complex analytical tasks.
Loading preview...
Model Overview
UwU-72B-Preview is a 72.7 billion parameter experimental research model developed by gghfez. This model was trained using a synthetic dataset generated by Qwen/QwQ-32B-Preview to investigate the potential of rank training in achieving comparable performance. It supports a substantial context length of 32768 tokens.
Key Capabilities
- Experimental Research: Focuses on exploring training methodologies, specifically rank training with synthetic data.
- Complex Reasoning: Utilizes a Chain-of-Thought (CoT) system prompt, encouraging step-by-step thinking for detailed problem-solving.
- High Context Length: Benefits from a 32768 token context window, allowing for processing and generating longer, more coherent responses.
- ChatML Format: Compatible with the ChatML chat template for structured conversational interactions.
Good For
- Researchers interested in synthetic data training and rank training techniques.
- Applications requiring detailed, step-by-step reasoning and analysis.
- Use cases benefiting from a large context window for extended dialogues or document processing.