Shuttle-3
Property | Value |
---|---|
Parameter Count | 72.7B |
Base Model | Qwen2.5-72B-Instruct |
License | Qwen License |
Training Data | 130M tokens |
Training Infrastructure | 4 A100 PCIe GPUs |
What is Shuttle-3?
Shuttle-3 is a state-of-the-art language model developed by ShuttleAI Inc., built upon the Qwen2.5-72B-Instruct architecture. This model represents a significant advancement in multilingual AI communication, specifically designed to emulate the writing quality of Claude 3 models while incorporating extensive role-playing capabilities.
Implementation Details
The model underwent intensive training on 130 million tokens over a 12-hour period using 4 A100 PCIe GPUs. It implements the ChatML format for prompting, ensuring structured and consistent interactions. The model uses BF16 tensor type for optimal performance and memory efficiency.
- Multilingual and code-focused pretraining
- Claude 3-style prose quality emulation
- Extensive role-play optimization
- ChatML-based interaction format
Core Capabilities
- Advanced multilingual communication
- Complex reasoning tasks
- Role-playing scenarios
- Chat-based interactions
- Text generation with high coherence
Frequently Asked Questions
Q: What makes this model unique?
Shuttle-3 stands out for its combination of large-scale parameters (72.7B) and specific optimization for role-playing scenarios while maintaining Claude 3-like writing quality. Its multilingual capabilities and efficient implementation make it versatile for various applications.
Q: What are the recommended use cases?
The model is particularly well-suited for complex chat applications, multilingual communication tasks, role-playing scenarios, and situations requiring nuanced reasoning. It's designed to handle both technical and creative writing tasks effectively.