AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Thinking Machines has released Inkling, a massive 975-billion-parameter multimodal AI model, on Hugging Face. While it offers advanced capabilities across text, images, and audio, its hardware requirements and lack of independent benchmarks limit immediate practical use.

Thinking Machines has released Inkling, a 975-billion-parameter multimodal AI model, on Hugging Face. The model is designed to process text, images, and audio within a one-million-token context window, making it one of the largest open models of its kind. This release is notable because it offers broad access to a model with unprecedented scale, though hardware demands are extremely high.

The Inkling model is a decoder-only Mixture-of-Experts architecture with 975 billion total parameters and 41 billion active during processing, trained on 45 trillion tokens across multiple data types. It supports multimodal inputs, including images and audio, which are converted into embeddings for processing alongside text. This multimodal capability is discussed in detail in the original analysis. The architecture uses 256 experts, with a combination of global and sliding-window attention, and features hierarchical image patching and mel-spectrogram audio representations.

Hugging Face reports that the BF16 checkpoint requires approximately 2 TB of VRAM, while the NVFP4 checkpoint needs about 600 GB, making full deployment feasible only on high-end hardware. The release includes support in several inference frameworks, such as Transformers, SGLang, vLLM, and llama.cpp, with options for hosted inference services. The model is positioned for domain-specific fine-tuning, especially for scientific, media, and enterprise applications involving mixed data types.

At a glance
announcementWhen: announced July 2026
The developmentThinking Machines has made Inkling, a large multimodal AI model, available on Hugging Face, marking a significant step in open access to high-scale AI models.

Implications of Inkling’s Release for AI Development

The release of Inkling represents a development in making large-scale multimodal AI models accessible, which could support research and application development in various fields such as scientific analysis, media, and enterprise data processing. However, the hardware requirements and absence of independent benchmarks currently limit its widespread adoption, especially among smaller teams or individual developers.

Its open availability may encourage further innovation but also raises considerations regarding safety, licensing, and transparency, which are yet to be clarified. The model’s capacity to handle complex multimodal data at scale could influence future AI applications, provided the technical barriers are addressed.

Amazon

high VRAM graphics card for AI training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Large Multimodal AI Models

Prior to Inkling, most large-scale multimodal models have been proprietary or limited in scope due to computational demands. The trend toward open models has been driven by organizations like OpenAI and Meta, but few have approached the scale of Inkling. The model’s training on 45 trillion tokens and its architecture reflect ongoing advances in sparse Mixture-of-Experts designs, which aim to balance scale and efficiency.

While earlier models like GPT-4 and multimodal variants such as GPT-4 Vision have demonstrated multimodal capabilities, they are typically not openly available at this scale. The release of Inkling on Hugging Face marks a shift toward greater openness, though the high hardware requirements continue to restrict direct usage to well-resourced entities.

“This model is large in scale.”

— Hugging Face

Amazon

professional GPU for large AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unverified Aspects of Inkling’s Performance and Licensing

Independent benchmark results, safety evaluations, and detailed licensing terms for Inkling are not yet available. It remains unclear how well the model performs across diverse multimodal tasks, especially in real-world scenarios, or how effectively the predictive layers improve inference speed in practice. The capabilities for video processing have also not been evaluated, and the licensing details are unspecified, raising questions about usage restrictions.

Amazon

AI inference server hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Testing and Validation of Inkling

Early testing by developers using Hugging Face’s supported inference frameworks will provide insights into the model’s real-world performance, latency, and accuracy. Independent researchers and organizations are expected to evaluate safety, benchmark its multimodal reasoning, and explore domain-specific fine-tuning. Clarifications around licensing, deployment costs, and hardware requirements are anticipated in the coming months, alongside potential updates from Thinking Machines and Hugging Face.

Amazon

multimodal AI model deployment hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Inkling and what makes it significant?

Inkling is a multimodal AI model with 975 billion parameters, capable of processing text, images, and audio simultaneously. Its large scale and open availability represent a notable development in AI research, although high hardware demands limit immediate practical use.

Can I run Inkling on a personal computer?

No. The model’s checkpoints require approximately 2 TB of VRAM for BF16 or 600 GB for NVFP4, making it inaccessible for typical consumer hardware. Hosted inference or specialized hardware is necessary.

Will Inkling process videos?

While the architecture supports image inputs with a temporal dimension, native video processing has not been evaluated or confirmed, so its video capabilities remain uncertain.

What are the licensing terms for Inkling?

The release describes Inkling as an open model but does not specify licensing restrictions, usage rights, or whether training code and data are publicly available. Details are expected to clarify in future updates.

When will independent benchmarks or safety evaluations be available?

These assessments are not yet available. Expect evaluations from early testers and third-party organizations over the next few months, which will help define the model’s strengths and limitations.

Source: ThorstenMeyerAI.com

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

When One Agent Isn’t Enough: Claude Now Builds Its Own Team of Agents on the Fly

Anthropic’s Claude now builds its own team of agents dynamically for complex tasks, enhancing performance on high-value projects.

Unlocking Frontier AI On A 512GB Mac Studio: What You Need To Know

Apple’s new Mac Studio with 512GB memory enables local running of frontier-scale AI models. Here’s what is confirmed, what remains unclear, and why it matters.

Is DeepSeek Threatening Anthropic’s Claude? Here’s What We Know

DeepSeek publicly announced its effort to compete with Anthropic’s Claude Code, signaling increased competition in AI-assisted software development tools.

Show HN: Run An 80B Qwen In 4.3 GB Of RAM On A Mac, And A 35B On An iPhone

A developer demonstrates running an 80-billion-parameter Qwen model on a Mac with 4.3 GB RAM and a 35B version on an iPhone, showcasing advanced optimization.