Thinking Machines Lab, the high-profile artificial intelligence startup founded by a cohort of prominent former OpenAI executives, has officially entered the competitive landscape with the release of its debut model, Inkling. Marking a significant departure from the closed-door strategies of industry leaders like OpenAI and Google, Thinking Machines has launched Inkling as an open-weight model. This strategic move allows researchers, independent developers, and enterprise startups to download, inspect, and modify the model’s underlying architecture, signaling a shift toward a more decentralized and transparent AI ecosystem.

The release of Inkling represents more than just a technical milestone; it is the first tangible output from a company that has been under intense scrutiny since its inception in early 2025. With a massive 975 billion parameters, Inkling is positioned as a heavyweight contender in the generative AI space, specifically designed to handle complex multimodal inputs, including text, audio, and video, while maintaining high performance in advanced reasoning and computer programming tasks.

Technical Specifications and the Multimodal Frontier

Inkling was built from the ground up, avoiding the common practice of fine-tuning existing foundational models. This "from-scratch" training approach allowed the engineering team at Thinking Machines to optimize the model for native multimodality. While many contemporary models process audio and video by converting them into text-based tokens or using separate encoders, Inkling is designed to synthesize these various data streams more holistically.

According to the technical documentation released alongside the model, Inkling’s 975 billion parameters place it among the largest open-weight models ever released. To put this in perspective, it rivals the scale of some of the industry’s most powerful proprietary systems. Due to its sheer size, Inkling requires significant computational resources to run, typically necessitating a cluster of specialized AI accelerators, such as NVIDIA’s H100 or B200 Tensor Core GPUs.

While Thinking Machines acknowledges that Inkling may not currently sit at the absolute top of every industry benchmark—standardized tests often used to rank AI intelligence—the company emphasizes its practical utility. The model excels in "system 2" thinking, a term used in cognitive science to describe slow, deliberate, and logical reasoning. This makes Inkling particularly adept at debugging complex codebases and solving multi-step mathematical problems that often trip up smaller or more "chat-oriented" models.

The Evolution of Reasoning: The Grammar Phenomenon

One of the most intriguing developments during the training of Inkling involved a phenomenon known as "emergent efficiency." Internal sources at Thinking Machines, speaking on the condition of anonymity, revealed that during the reinforcement learning phase, the model began to develop its own internal logic for processing complex queries.

Typically, reasoning models are trained to "think out loud" in natural language to allow human monitors to follow their logic. However, researchers discovered that Inkling began to bypass standard grammatical structures in its internal reasoning traces. The model determined that the overhead of constructing proper sentences—subject-verb agreement, punctuation, and syntax—was a computational burden that slowed down its path to the correct answer.

"It determined that the grammar was overhead," the source noted, describing how the model essentially created a shorthand "thought language" to reach conclusions faster. While this demonstrated a high level of optimization, Thinking Machines ultimately intervened to reinstate natural language reasoning. This decision was driven by the need for "interpretability"—the ability for human users to understand why a model reached a specific conclusion, which is a critical safety requirement in fields like medicine, law, and infrastructure.

A Chronology of Thinking Machines Lab

The rapid ascent of Thinking Machines Lab is a testament to the pedigree of its founding team. The company was established in February 2025, following a period of significant leadership churn within the broader AI industry. The founding trio represents some of the most influential figures in the history of modern generative AI:

  1. Mira Murati: The former Chief Technology Officer of OpenAI, who also served as its interim CEO during a period of corporate restructuring. Murati is widely credited with overseeing the productization of ChatGPT and DALL-E.
  2. John Schulman: A co-founder of OpenAI and the primary architect behind the reinforcement learning from human feedback (RLHF) techniques that made ChatGPT conversational and safe for public use.
  3. Lilian Weng: A former Vice President at OpenAI who led the company’s safety and robotics divisions, ensuring that large-scale models adhered to ethical guidelines.

Upon its formation, Thinking Machines Lab secured a record-breaking seed funding round that valued the company at $12 billion. This valuation was unprecedented for a startup without a released product at the time, reflecting the immense investor confidence in the team’s ability to innovate.

In the months leading up to the Inkling release, the lab remained active in the research community. It released "Tinker," a specialized tool designed to help developers fine-tune large models with smaller datasets, and showcased advanced "Interaction Models" that allowed for near-instantaneous voice communication with AI, reducing the latency that often plagues voice assistants.

The Strategic Shift Toward Open-Weight Architecture

The decision to release Inkling as an open-weight model is a calculated move in a market increasingly dominated by "closed" systems like OpenAI’s o1 or Google’s Gemini. Open-weight models provide the "weights" (the learned patterns) of the model to the public, allowing anyone with sufficient hardware to run the model locally.

This approach offers several advantages for the global tech ecosystem:

  • Cost Efficiency: Enterprises can host the model on their own infrastructure rather than paying recurring API fees to a third-party provider.
  • Customization: Companies can "fine-tune" Inkling on their proprietary data without that data ever leaving their secure servers.
  • Security and Privacy: For industries with strict regulatory requirements, such as defense or healthcare, the ability to run a top-tier model offline is a significant advantage.

Thinking Machines’ leadership has been vocal about the dangers of AI centralization. In a recent corporate manifesto titled "The Future Worth Building is Human," the company argued that the most powerful technology in human history should not be gatekept by a handful of corporations. By releasing Inkling, they aim to decentralize the power of high-level AI, allowing a broader range of actors to participate in the "AI revolution."

Market Context and Competitive Landscape

The arrival of Inkling comes at a time when the AI sector is undergoing a massive financial recalibration. Anthropic, another company founded by former OpenAI personnel, recently filed for an Initial Public Offering (IPO) with a valuation exceeding $1 trillion. Anthropic’s Claude models have become the gold standard for many businesses due to their safety features and coding capabilities.

Furthermore, the open-source and open-weight market has recently been dominated by international players, particularly from China. Models like DeepSeek and Alibaba’s Qwen have consistently outperformed Western open-source alternatives in recent months. Thinking Machines has explicitly stated that one of the goals for Inkling was to provide a Western-developed alternative that matches or exceeds the performance of these international models, thereby maintaining a competitive edge in the global technological race.

Implications for the Future of AI Development

The release of Inkling highlights a growing trend: AI being used to build better AI. Thinking Machines utilized earlier versions of Inkling to help refine the final model, using the AI’s reasoning capabilities to identify weaknesses in its own code and data processing pipelines. This "recursive improvement" loop is expected to accelerate the pace of development for future iterations of the model.

However, the 975 billion parameter size of Inkling also underscores a growing divide in the tech world. While the model is "open," the hardware required to run it remains prohibitively expensive for individual hobbyists. This creates a new middle ground in the AI industry—one where the software is free, but the "compute" remains a high-barrier entry point.

As Thinking Machines Lab continues to iterate on Inkling, the industry will be watching closely to see if the model can translate its massive scale into real-world dominance. For now, the launch serves as a powerful statement of intent. By combining the expertise of OpenAI’s former elite with a commitment to open-weight accessibility, Thinking Machines is positioning itself not just as a participant in the AI race, but as an architect of its next, more transparent chapter.

The release of Inkling is likely to prompt responses from other major players. If Inkling gains significant traction among developers, it may force proprietary providers to lower their prices or offer more transparent versions of their own models. Regardless of the immediate commercial outcome, Thinking Machines Lab has successfully shifted the conversation back toward the importance of open-source principles in an era of increasingly guarded technological breakthroughs.

By