Dynamical-Systems/Dynamical-SDL1-35B-A3B

TEXT GENERATIONPricing:Input $0.4 / Cached $0.07 / Output $4Concurrent Unit Cost:3Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 6, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Dynamical-Systems/Dynamical-SDL1-35B-A3B is a 35.1 billion parameter Qwen3.6 mixture-of-experts model with 3 billion active parameters, post-trained for scientific decision-making in experimental campaigns. It functions as an open-weights policy to guide evidence acquisition, belief updates, and experiment selection within partially observable decision environments. This model excels at replaying recorded physical campaigns and optimizing scientific agency under a fixed budget, demonstrating robust performance in the SDL-1 benchmark environment. Its primary use is as a research artifact for experimental-campaign replay, evaluation, and method development.

Loading preview...

Model Overview

Dynamical-SDL1-35B-A3B is a 35.1 billion total parameter, 3 billion active parameter Qwen3.6 mixture-of-experts model developed by Dynamical-Systems. It was initialized from Qwen/Qwen3.6-35B-A3B and specifically post-trained within the SDL-1 environment to act as an open-weights policy for scientific decision-making. The model's core function is to navigate partially observable decision environments, making choices on evidence acquisition, belief revision, and subsequent experiment execution under budget constraints.

Key Capabilities

  • Scientific Decision Policy: Guides experimental campaigns by deciding what evidence to acquire, how beliefs should change, and which experiment to run next.
  • Experimental Campaign Replay: Optimized for replaying recorded physical campaigns within the SDL-1 environment.
  • Robust Performance: Achieved a negative held-out log loss of -1.4664 in the SDL-1 benchmark, completing all 12 campaign branches without overclaiming.
  • Inference System: Utilizes a comprehensive inference package including a frozen system prompt, behavior rules, and a three-draw forecast pooling method for enhanced decision quality.
  • Outcome-Grounded RL: Trained using outcome-grounded reinforcement learning to assimilate evidence and scope claims against evaluator-private recorded outcomes.

Good For

  • Research and Development: Ideal as a research artifact for developing and evaluating methods in experimental-campaign replay and scientific agency.
  • Simulated Scientific Environments: Suitable for use with environments that separate evaluator-private outcomes from policy-visible evidence and validate structured actions.
  • Understanding Scientific Process Automation: Provides a framework for exploring automated decision-making in complex scientific discovery processes.