RHYu2233/LaCT-Motion-SFT

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 17, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

RHYu2233/LaCT-Motion-SFT is a 3.1 billion parameter curriculum supervised fine-tuned checkpoint of LaCT-Motion, developed by RHYu2233, for text-to-motion generation. Built upon Qwen2.5-3B-Instruct, this model extends its vocabulary with 512 motion codes and latent-reasoning special tokens, enabling reinforced latent planning. It specializes in generating human motion sequences from text descriptions, leveraging a latent chain-of-thought approach. This model is primarily intended for non-commercial research and evaluation in text-to-motion synthesis.

Loading preview...

LaCT-Motion-SFT: Text-to-Motion Generation Checkpoint

This model, RHYu2233/LaCT-Motion-SFT, is a 3.1 billion parameter supervised fine-tuned (SFT) checkpoint of LaCT-Motion (Latent Chain-of-Thought for Motion). It is based on the Qwen2.5-3B-Instruct architecture and is designed for advanced text-to-motion generation, as detailed in an ECCV 2026 paper. The model incorporates a unique approach of reinforced latent planning, utilizing an extended vocabulary that includes 512 motion codes and specific latent-reasoning tokens.

Key Capabilities & Features

  • Text-to-Motion Generation: Translates textual descriptions into human motion sequences.
  • Latent Chain-of-Thought: Employs a novel latent planning mechanism for motion synthesis, with settings like c_thought: 2 and max_latent_stage: 8.
  • Specialized Vocabulary: Extends the base Qwen vocabulary with 512 motion codes (<Motion_0> to <Motion_511>), motion delimiters, and latent-reasoning tokens.
  • Base Model: Fine-tuned from Qwen/Qwen2.5-3B-Instruct, inheriting its 3.09 billion parameters and 32768 token context length.
  • Training Data: Trained on processed splits of the HumanML3D dataset.
  • Precision: Distributed in bfloat16 safetensors.

Use Cases & Licensing

This checkpoint is specifically designed to initialize Stage 2 GRPO runs for refining fully latent policies in text-to-motion tasks. It can be used for evaluating motion generation on HumanML3D or for generating motions via the provided demo scripts. The model is released under the Qwen RESEARCH LICENSE AGREEMENT, restricting its use to non-commercial research and evaluation only. Commercial applications require a license from Alibaba Cloud.