amazon/ALoDLM-8B
ALoDLM-8B is an 8 billion parameter diffusion language model developed by Amazon, initialized from the Qwen3 backbone, with a 32768 token context length. It generates text and code by adaptively applying shared transformer layers, preserving and refining unresolved tokens' latent states. This model excels in mathematical reasoning and code generation by dynamically allocating computation, offering improved efficiency and accuracy for these tasks.
Loading preview...
ALoDLM-8B: Adaptively Looped Diffusion Language Model
ALoDLM-8B is an 8 billion parameter diffusion language model developed by Amazon, built upon the Qwen3 backbone. This model introduces a novel approach to text and code generation by adaptively looping shared transformer layers. Unlike traditional fixed-depth denoisers, ALoDLM-8B preserves and refines the latent states of unresolved tokens through additional recurrent passes, while committed tokens provide context. This adaptive computation allocation is controlled by predictive confidence and a learned halting policy, allowing for dynamic computational depth based on token resolution.
Key Capabilities
- Adaptive Computation Allocation: Dynamically allocates computational resources by refining unresolved token representations across recurrent passes, leading to improved efficiency.
- Enhanced Reasoning and Code Generation: Trained on extensive datasets including mathematical problems, worked solutions, programming tasks, and instruction-formatted text, making it highly proficient in these domains.
- Persistent Latent Refinement: Reuses shared layers to continuously refine hidden states of tokens that require more processing.
- Inference-time Control: Offers control over the trade-off between recurrent refinement and parallel token commitment through adjustable gate (
q) and entropy (τ) thresholds.
Good For
- Noncommercial Research: Ideal for research into diffusion language modeling, adaptive computation, and advanced text generation techniques.
- Mathematical Reasoning Tasks: Demonstrates strong performance on benchmarks like GSM8K and MATH-500, often outperforming other diffusion and autoregressive models in its class.
- Code Generation: Excels in programming tasks, achieving high scores on benchmarks such as MBPP and HumanEval, making it suitable for generating and refining code.
- Exploring Computational Efficiency: Useful for investigating the effects of changing inference-time computation and optimizing resource allocation in language models.