Ant Group Unveils Ling-3.0-Flash Delivering Top-Tier Performance at a Fraction of the Parameter Scale

27 Jul 2026
HANGZHOU, China

Ant Group today announced the release of Ling-3.0-Flash, a next-generation native hybrid-reasoning foundational model engineered specifically for production-grade AI agent workflows. Designed to deliver rapid response capabilities, it serves as a high-speed execution node that offers a superior balance of intelligence density and cost-efficiency.

This press release features multimedia. View the full release here: https://www.businesswire.com/news/home/20260726584441/en/

Ling-3.0-Flash delivers strong performance across multiple core benchmarks.

Ling-3.0-Flash delivers strong performance across multiple core benchmarks.

Featuring 124B total parameters with only 5.1B active parameters per token, Ling-3.0-Flash achieves remarkable performance despite its streamlined footprint. It matches or surpasses industry-leading models with two to three times its parameter scale across core benchmarks, including foundational reasoning, instruction following, and long-context processing.

Architectural Innovation for Efficiency

Ling-3.0-Flash moves away from the traditional approach of simply scaling parameter counts. Instead, it is built from the ground up with a native hybrid-linear attention architecture. By alternating KDA (Kimi Delta Attention) and MLA layers at a 5:1 ratio, the model optimally balances long-context efficiency with robust state memory.

Key architectural advancements include:

  • Upgraded KDA: Evolving from the previous Lightning Attention, KDA introduces fine-grained diagonal gating in Delta Rule state updates, allowing the model to retain critical information more precisely when processing lengthy documents and extensive codebases.
  • Optimized Mixture-of-Expert (MoE) Compute: The expert activation ratio per token has been compressed from 1/32 in the previous generation to 1/64, yielding a significantly higher "efficiency leverage."
  • Extended Context Window: The model natively supports a 256K context window and can seamlessly scale to 1M tokens.

Purpose-Built for Agent Workflows

Rather than aiming to replace ultra-large, general-purpose reasoning models, Ling-3.0-Flash is designed to complete the "planning-execution separation" paradigm in AI workflows. It serves as a cost-controllable, fast, and highly stable execution node, delegating deep planning and high-frequency execution to specialized models.

To support this, Ling-3.0-Flash has been deeply refined for real-world agent scenarios, expanding its training to over 10,000 interactive environments. It features enhanced self-correction and long-horizon planning mechanisms, enabling autonomous, end-to-end delivery in complex tasks such as coding, task decomposition, and deep multi-source research. This resolves common issues of deviation or context loss in traditional models during large-scale operations.

Engineering for Speed and Stability

To ensure fast and reliable agent performance, Ant Group has paired Ling-3.0-Flash with a supporting engineering and collaboration architecture:

  • Reduced Latency: A cluster-level hierarchical caching system eliminates redundant computations in long conversations and multi-turn interactions, reducing Time-to-First-Token (TTFT) for long inputs by 60% to over 80%.
  • Enhanced Stability: An upgraded multi-agent collaboration architecture enables different agents to divide labor and cross-validate outputs, significantly reducing the risk of misjudgments by a single model and providing robust support for high-frequency online services.

Ling-3.0-Flash is now available on OpenRouter and Vercel AI Gateway, offering a free API through August 3, 2026. Following this limited-time free access period, the model weights will be open-sourced to support further development and innovation within the global AI community.

Developers are encouraged to integrate Ling-3.0-Flash into their coding, search, research, and tool-use workflows to experience its high-speed execution and stable tool-calling capabilities.

About Ant Group

Ant Group is a global digital technology provider and the operator of Alipay, a leading internet services platform in China, connecting over one billion users to more than 10,000 types of consumer services from partners. Through innovative products and solutions powered by AI, blockchain and other technologies, Ant Group supports partners across industries to thrive through digital transformation in an ecosystem for inclusive and sustainable development. For more information, visit www.antgroup.com.

 

© Business Wire, Inc.

Haftungsausschluss :
Diese Pressemitteilung ist kein von AFP erstelltes Dokument. AFP übernimmt keine Verantwortung für ihren Inhalt. Bei Fragen wenden Sie sich bitte an die im Text der Pressemitteilung genannten Kontaktpersonen/Stellen.