📊 Full opportunity report: Revolutionize Your Tech With Inkling, The Latest In AI on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Thinking Machines has released Inkling, a massive 975-billion-parameter multimodal AI model, on Hugging Face. While it offers advanced capabilities across text, images, and audio, its hardware requirements and lack of independent benchmarks limit immediate practical use.

Thinking Machines has released Inkling, a 975-billion-parameter multimodal AI model, on Hugging Face. The model is designed to process text, images, and audio within a one-million-token context window, making it one of the largest open models of its kind. This release is notable because it offers broad access to a model with unprecedented scale, though hardware demands are extremely high.

The Inkling model is a decoder-only Mixture-of-Experts architecture with 975 billion total parameters and 41 billion active during processing, trained on 45 trillion tokens across multiple data types. It supports multimodal inputs, including images and audio, which are converted into embeddings for processing alongside text. This multimodal capability is discussed in detail in the original analysis. The architecture uses 256 experts, with a combination of global and sliding-window attention, and features hierarchical image patching and mel-spectrogram audio representations.

Hugging Face reports that the BF16 checkpoint requires approximately 2 TB of VRAM, while the NVFP4 checkpoint needs about 600 GB, making full deployment feasible only on high-end hardware. The release includes support in several inference frameworks, such as Transformers, SGLang, vLLM, and llama.cpp, with options for hosted inference services. The model is positioned for domain-specific fine-tuning, especially for scientific, media, and enterprise applications involving mixed data types.

At a glance
announcementWhen: announced July 2026
The developmentThinking Machines has made Inkling, a large multimodal AI model, available on Hugging Face, marking a significant step in open access to high-scale AI models.
At a glance
announcementWhen: announced on Hugging Face; the source m…
The developmentThinking Machines has made its Inkling multimodal model available through Hugging Face with day-one support from several major inference frameworks.

Implications of Inkling’s Release for AI Development

The release of Inkling represents a development in making large-scale multimodal AI models accessible, which could support research and application development in various fields such as scientific analysis, media, and enterprise data processing. However, the hardware requirements and absence of independent benchmarks currently limit its widespread adoption, especially among smaller teams or individual developers.

Its open availability may encourage further innovation but also raises considerations regarding safety, licensing, and transparency, which are yet to be clarified. The model’s capacity to handle complex multimodal data at scale could influence future AI applications, provided the technical barriers are addressed.

GIGABYTE Radeon™ AI PRO R9700 AI TOP 32G Graphics Card, Turbo Fan Cooling System, 32GB GDDR6, GV-R9700AI TOP-32GD Video Card

GIGABYTE Radeon™ AI PRO R9700 AI TOP 32G Graphics Card, Turbo Fan Cooling System, 32GB GDDR6, GV-R9700AI TOP-32GD Video Card

Powered by Radeon AI PRO R9700 – Supercharge you workflow with the cutting-edge RDNA 4 Architecture and 2nd-gen…

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Large Multimodal AI Models

Prior to Inkling, most large-scale multimodal models have been proprietary or limited in scope due to computational demands. The trend toward open models has been driven by organizations like OpenAI and Meta, but few have approached the scale of Inkling. The model’s training on 45 trillion tokens and its architecture reflect ongoing advances in sparse Mixture-of-Experts designs, which aim to balance scale and efficiency.

While earlier models like GPT-4 and multimodal variants such as GPT-4 Vision have demonstrated multimodal capabilities, they are typically not openly available at this scale. The release of Inkling on Hugging Face marks a shift toward greater openness, though the high hardware requirements continue to restrict direct usage to well-resourced entities.

“This model is large in scale.”

— Hugging Face

MX3 M.2 AI Accelerator

MX3 M.2 AI Accelerator

High-Performance AI Processing: The MX3 is designed to handle the most demanding AI computer vision workloads, delivering exceptional…

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unverified Aspects of Inkling’s Performance and Licensing

Independent benchmark results, safety evaluations, and detailed licensing terms for Inkling are not yet available. It remains unclear how well the model performs across diverse multimodal tasks, especially in real-world scenarios, or how effectively the predictive layers improve inference speed in practice. The capabilities for video processing have also not been evaluated, and the licensing details are unspecified, raising questions about usage restrictions.

GRADIO 5 WITH SERVER-SIDE RENDERING: INSTANT ML UIS AND AI PLAYGROUND: Build Production Interfaces with ChatInterface, Blocks API, and Natural Language App Generation

GRADIO 5 WITH SERVER-SIDE RENDERING: INSTANT ML UIS AND AI PLAYGROUND: Build Production Interfaces with ChatInterface, Blocks API, and Natural Language App Generation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Testing and Validation of Inkling

Early testing by developers using Hugging Face’s supported inference frameworks will provide insights into the model’s real-world performance, latency, and accuracy. Independent researchers and organizations are expected to evaluate safety, benchmark its multimodal reasoning, and explore domain-specific fine-tuning. Clarifications around licensing, deployment costs, and hardware requirements are anticipated in the coming months, alongside potential updates from Thinking Machines and Hugging Face.

Vvikizy Dual LGA 2011 E5 Server Motherboard, C602 Chipset Support for 8 DDR3 Slots 256GB RAM, with Multiple PCIe 3.0 Slots for AI Training GPU Workstation

Vvikizy Dual LGA 2011 E5 Server Motherboard, C602 Chipset Support for 8 DDR3 Slots 256GB RAM, with Multiple PCIe 3.0 Slots for AI Training GPU Workstation

[DUAL CPU POWERHOUSE FOR PROFESSIONAL WORKLOADS] This high performance workstation motherboard features dual LGA 2011 sockets supporting E5…

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Inkling and what makes it significant?

Inkling is a multimodal AI model with 975 billion parameters, capable of processing text, images, and audio simultaneously. Its large scale and open availability represent a notable development in AI research, although high hardware demands limit immediate practical use.

Can I run Inkling on a personal computer?

No. The model’s checkpoints require approximately 2 TB of VRAM for BF16 or 600 GB for NVFP4, making it inaccessible for typical consumer hardware. Hosted inference or specialized hardware is necessary.

Will Inkling process videos?

While the architecture supports image inputs with a temporal dimension, native video processing has not been evaluated or confirmed, so its video capabilities remain uncertain.

What are the licensing terms for Inkling?

The release describes Inkling as an open model but does not specify licensing restrictions, usage rights, or whether training code and data are publicly available. Details are expected to clarify in future updates.

When will independent benchmarks or safety evaluations be available?

These assessments are not yet available. Expect evaluations from early testers and third-party organizations over the next few months, which will help define the model’s strengths and limitations.

Source: ThorstenMeyerAI.com

You May Also Like

The Door: Why the Interface Is Worth More Than the Model

SpaceX’s $60B purchase of a coding interface highlights the growing importance of interface ownership over AI models in distribution and control.

The Trust Shock: What Suspending Fable 5 Means for US AI, Its Rivals, and the World

The US government’s sudden halt of Anthropic’s Fable 5 raises questions about AI trust, US dominance, and industry stability amid new export controls.

Evolution of Artificial Intelligence videos just in 4 years is mind blowing

The rapid progression of AI-generated videos in just four years showcases unprecedented advancements in technology, transforming content creation and perception.

The Classroom Tech Flops That Could Guide Ai’s Next Evolution

Failed classroom tech reveals crucial lessons that could shape AI’s future—discover what went wrong and how it can lead to smarter innovations.