The Secret Weapon In AI: Inkling By Thinking Machines

📊 Full opportunity report: The Secret Weapon In AI: Inkling By Thinking Machines on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Thinking Machines has unveiled Inkling, a 975-billion-parameter multimodal AI model, on Hugging Face. Its open availability signals a significant step in large-scale AI, as detailed in the original analysis, though hardware and licensing details remain uncertain.

Thinking Machines has officially released Inkling, a 975-billion-parameter multimodal model, on Hugging Face. The release makes it accessible to developers and researchers, despite its demanding hardware requirements, marking a notable advancement in open large-scale AI models.

Inkling is described as a decoder-only Mixture-of-Experts model with 975 billion total parameters and 41 billion active during processing, trained on 45 trillion tokens across text, images, audio, and video. Its architecture employs 256 experts, using a combination of global and sliding-window attention mechanisms, with image inputs processed through hierarchical patching and audio converted into mel-spectrograms.

The model is available via Hugging Face with support in popular inference tools, including Transformers and llama.cpp, and can be accessed through hosted inference providers. However, running the full model requires extensive hardware—approximately 2 TB of VRAM for BF16 checkpoints and 600 GB for NVFP4—limiting direct deployment to specialized hardware setups. The release lacks independent benchmark results, safety evaluations, or detailed licensing information, leaving many performance and usage questions open.

At a glance
breakingWhen: announced July 2026
The developmentThinking Machines has released Inkling, a large multimodal model, on Hugging Face, marking a major development in open AI models with high hardware requirements.
At a glance
announcementWhen: announced on Hugging Face; the source m…
The developmentThinking Machines has made its Inkling multimodal model available through Hugging Face with day-one support from several major inference frameworks.

Implications of Inkling’s Open Multimodal Scale

The release of Inkling represents a significant step toward more accessible large-scale AI models that can process multiple data types within a single framework. Its open availability could accelerate research and development in fields like scientific analysis, media processing, and enterprise applications, especially for tasks requiring integrated text, image, and audio reasoning.

However, the substantial hardware requirements pose a barrier to widespread adoption, limiting full deployment to organizations with advanced infrastructure. The lack of independent benchmarks and licensing clarity also raises questions about its safety, reliability, and practical use in production environments.

GIGABYTE Radeon™ AI PRO R9700 AI TOP 32G Graphics Card, Turbo Fan Cooling System, 32GB GDDR6, GV-R9700AI TOP-32GD Video Card

GIGABYTE Radeon™ AI PRO R9700 AI TOP 32G Graphics Card, Turbo Fan Cooling System, 32GB GDDR6, GV-R9700AI TOP-32GD Video Card

Powered by Radeon AI PRO R9700 – Supercharge you workflow with the cutting-edge RDNA 4 Architecture and 2nd-gen…

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of Large Multimodal AI Development

Large multimodal models have been emerging as a key frontier in AI research, with previous efforts like GPT-4, PaLM-E, and Flamingo demonstrating capabilities in integrating text, images, and videos. However, most of these models remain proprietary or limited in scale. Inkling’s release by Thinking Machines, a company known for pushing the boundaries of model size, underscores a trend toward open, high-parameter models designed to handle complex, real-world data.

Prior to Inkling, the industry has grappled with balancing model scale, accessibility, and hardware demands. The move to make such a large model openly available marks a notable shift, though practical deployment remains constrained by infrastructure needs and the lack of comprehensive independent evaluations.

“This model is huge.”

— Hugging Face spokesperson

Cerebras GPT: Wafer-Scale Architectures for Large Language Models

Cerebras GPT: Wafer-Scale Architectures for Large Language Models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unverified Aspects and Performance Uncertainties

Independent benchmark results, safety assessments, and detailed licensing terms are not yet available. The actual performance of Inkling on practical workloads, especially in video processing and multimodal reasoning, remains unverified. It is also unclear how well the model handles real-world tasks or how effective the speculative prediction layers are in production environments.

Multimodal AI Systems Engineering: Building Production Vision-Language Models, Document AI, and Cross-Modal Retrieval Pipelines (Production AI Engineering Series)

Multimodal AI Systems Engineering: Building Production Vision-Language Models, Document AI, and Cross-Modal Retrieval Pipelines (Production AI Engineering Series)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Testing and Evaluation

Developers and organizations will begin testing Inkling through supported inference engines, focusing on latency, accuracy, and hardware requirements. Expect upcoming independent evaluations, benchmark disclosures, and safety assessments to clarify its capabilities and limitations. Further, licensing details and potential fine-tuning applications are likely to be announced in the near future.

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Inkling?

Inkling is a large, open multimodal AI model from Thinking Machines, capable of processing text, images, and audio, with 975 billion parameters.

Can Inkling process videos?

While the architecture supports image inputs with a temporal dimension, native video processing performance has not been evaluated, and no definitive claims about video capabilities are available yet.

Is Inkling suitable for running on personal computers?

Likely not. The model requires approximately 2 TB of VRAM for full deployment, making it impractical for typical consumer hardware. Hosted inference or specialized infrastructure is recommended.

What are the licensing and safety considerations?

Details about licensing, usage restrictions, safety evaluations, and training data are not yet disclosed, leaving questions about its commercial and research deployment open.

When will more performance data be available?

Further evaluations, benchmarks, and safety assessments are expected as developers and third parties test Inkling in real-world scenarios over the coming months.

Source: ThorstenMeyerAI.com

You May Also Like

ECC And DDR5

New developments confirm ECC support for DDR5 RAM, impacting server and high-reliability computing markets. Details on compatibility and availability are emerging.

Today’s Wordle Hints, Answer and Help for June 19, #1826

Get the confirmed Wordle answer, hints, and help for June 19, #1826. Find out what today’s puzzle is, why it matters, and what remains uncertain.

The SSD Squeeze: Why Storage Joined the Party

Enterprise and consumer SSD prices soar as AI drives unprecedented storage demand and supply constraints tighten, impacting the entire market.

The $60 Billion Bargain: Why Cursor Could Be a Steal for SpaceX

SpaceX’s planned $60B all-stock purchase of Cursor maker Anysphere is signed but not closed, with growth, margins and review risks in focus.