📊 Full opportunity report: Meet Inkling: The AI Breakthrough You Need To Know About on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Thinking Machines has released Inkling, a 975-billion-parameter multimodal AI model, on Hugging Face. For more details, see the original analysis. Its open availability and scale are notable, but hardware requirements and evaluation details remain unclear.

Thinking Machines has released Inkling, a 975-billion-parameter multimodal AI model, on Hugging Face. The model is designed to process text, images, and audio within a claimed one-million-token context window, marking a significant development in large-scale AI. This reflects the advancements discussed in recent industry coverage. Its open access makes it notable for researchers and developers, though the hardware demands are substantial.

Inkling is described as a decoder-only Mixture-of-Experts model with 975 billion total parameters, of which 41 billion are active during processing. It was trained on 45 trillion tokens across multiple modalities, including text, images, audio, and video. The architecture employs 256 experts, using a combination of global and sliding-window attention, along with other techniques aimed at optimizing local and global processing.

The model supports multimodal inputs by converting images into hierarchical patches and audio into mel-spectrograms. It includes checkpoints optimized for lower-precision inference, such as BF16 and NVFP4, but these require extensive VRAM—around 2 TB for BF16 and 600 GB for NVFP4—making full deployment impractical for most individual users. Instead, many will rely on hosted inference services or quantized versions.

Hugging Face reports initial support for Inkling in popular frameworks like Transformers, SGLang, vLLM, and llama.cpp, facilitating integration and testing. Learn more about the capabilities in the original analysis. The model is positioned for domain-specific fine-tuning, particularly in scientific, media analysis, and enterprise applications involving mixed data types.

At a glance
announcementWhen: announced July 2026
The developmentThinking Machines has made Inkling, a large-scale multimodal AI model, publicly available on Hugging Face, sparking interest and questions about its capabilities and limitations.
At a glance
announcementWhen: announced on Hugging Face; the source m…
The developmentThinking Machines has made its Inkling multimodal model available through Hugging Face with day-one support from several major inference frameworks.

Implications of Inkling’s Scale and Accessibility

The release of Inkling represents a notable development in the field of multimodal AI, combining a large number of parameters with open access. While its scale allows for advanced reasoning across text, images, and audio, the hardware requirements limit direct deployment to specialized organizations. The open availability could support research and application development, but the absence of independent benchmarks and licensing details raises questions about its practical performance and safety.

For industries such as scientific research, media analysis, and enterprise data processing, Inkling could serve as a tool for integrated multimodal reasoning. However, its high computational demands mean most users will depend on cloud-based services or future optimized versions.

Amazon

high VRAM graphics card for AI training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Large-Scale Multimodal Models

Prior to Inkling, large multimodal models have been developed by several organizations, but few have matched its scale of 975 billion parameters. The trend toward open models has increased, with some models providing limited access, but none at this scale with such multimodal versatility. Training on 45 trillion tokens across diverse data types reflects a push toward more general-purpose AI systems capable of understanding and reasoning across multiple modalities simultaneously.

Hardware constraints have historically limited the deployment of such models, with only major research labs and large corporations able to operate them directly. The release of Inkling, with its open access and support for multiple input types, indicates a shift toward broader experimentation, albeit with significant technical barriers for most users.

“This model is large.”

— Hugging Face’s Inkling release article

Amazon

multimodal AI development tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unverified Aspects and Performance Expectations

Independent benchmark results, safety evaluations, and detailed licensing information for Inkling have not yet been released. It remains unclear how the model performs across diverse workloads, especially in real-world scenarios involving video and complex multimodal reasoning. The actual speed, accuracy, and safety of the model in practical applications are still unknown, pending third-party testing and validation.

Amazon

large-scale AI inference servers

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Expected Developments and Testing Milestones

Developers and research organizations will begin testing Inkling through supported frameworks, aiming to evaluate its performance, latency, and safety. Independent evaluations, benchmark disclosures, and safety assessments are anticipated in the coming months. Additionally, efforts to develop domain-specific fine-tuned versions and explore practical deployment options are likely to follow.

Amazon

AI model deployment hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Inkling and why is it significant?

Inkling is a 975-billion-parameter multimodal AI model from Thinking Machines, capable of processing text, images, and audio. Its open release and scale mark a significant development in AI research and application potential.

Can Inkling process videos or real-time multimedia?

While the architecture supports inputs with temporal dimensions, native video processing performance has not been evaluated, and no definitive claims about video capabilities have been made.

Is Inkling available for individual use on personal computers?

No. The hardware requirements are extremely high, with estimates of 2 TB of VRAM for full BF16 deployment. Most users will rely on hosted inference or optimized, quantized versions.

What are the licensing and safety considerations?

The licensing terms, safety evaluations, and detailed training data disclosures have not been publicly provided, leaving questions about usage restrictions and safety open for now.

What are the next steps for researchers interested in Inkling?

They can begin testing the model via supported frameworks, await independent benchmark results, and monitor updates from Hugging Face and Thinking Machines for safety and licensing disclosures.

Source: ThorstenMeyerAI.com

You May Also Like

Today’s NYT Connections Hints, Answers and Help for July 1, #1116

Get the latest hints, answers, and tips for today’s NYT Connections puzzle (#1116) released on July 1, 2024. Find out what is confirmed and what remains unclear.

How Do Satellite Phones Work (and How Are They Different)?

The intriguing way satellite phones connect via orbiting satellites offers unique benefits and differences from regular phones, and you’ll want to learn more.

GTA 6: Price, release date, pre-orders and everything else you need to know

Official details on GTA 6’s release date, pricing, pre-order options, and what is still unknown, based on latest reports and leaks.

Watch an AI-Driven Company Fight for Survival in Real Time

Watch an AI-managed startup struggle live in public, facing crises, making decisions, and risking its survival — all while demonstrating the importance of discipline and trust.