The Secret Weapon In AI: Inkling By Thinking Machines
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Thinking Machines has unveiled Inkling, a 975-billion-parameter multimodal AI model, on Hugging Face. Its open availability signals a significant step in large-scale AI, as detailed in the original analysis, though hardware and licensing details remain uncertain.

Thinking Machines has officially released Inkling, a 975-billion-parameter multimodal model, on Hugging Face. The release makes it accessible to developers and researchers, despite its demanding hardware requirements, marking a notable advancement in open large-scale AI models.

Inkling is described as a decoder-only Mixture-of-Experts model with 975 billion total parameters and 41 billion active during processing, trained on 45 trillion tokens across text, images, audio, and video. Its architecture employs 256 experts, using a combination of global and sliding-window attention mechanisms, with image inputs processed through hierarchical patching and audio converted into mel-spectrograms.

The model is available via Hugging Face with support in popular inference tools, including Transformers and llama.cpp, and can be accessed through hosted inference providers. However, running the full model requires extensive hardware—approximately 2 TB of VRAM for BF16 checkpoints and 600 GB for NVFP4—limiting direct deployment to specialized hardware setups. The release lacks independent benchmark results, safety evaluations, or detailed licensing information, leaving many performance and usage questions open.

At a glance
breakingWhen: announced July 2026
The developmentThinking Machines has released Inkling, a large multimodal model, on Hugging Face, marking a major development in open AI models with high hardware requirements.

Implications of Inkling’s Open Multimodal Scale

The release of Inkling represents a significant step toward more accessible large-scale AI models that can process multiple data types within a single framework. Its open availability could accelerate research and development in fields like scientific analysis, media processing, and enterprise applications, especially for tasks requiring integrated text, image, and audio reasoning.

However, the substantial hardware requirements pose a barrier to widespread adoption, limiting full deployment to organizations with advanced infrastructure. The lack of independent benchmarks and licensing clarity also raises questions about its safety, reliability, and practical use in production environments.

Amazon

high VRAM GPU for AI inference

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of Large Multimodal AI Development

Large multimodal models have been emerging as a key frontier in AI research, with previous efforts like GPT-4, PaLM-E, and Flamingo demonstrating capabilities in integrating text, images, and videos. However, most of these models remain proprietary or limited in scale. Inkling’s release by Thinking Machines, a company known for pushing the boundaries of model size, underscores a trend toward open, high-parameter models designed to handle complex, real-world data.

Prior to Inkling, the industry has grappled with balancing model scale, accessibility, and hardware demands. The move to make such a large model openly available marks a notable shift, though practical deployment remains constrained by infrastructure needs and the lack of comprehensive independent evaluations.

“This model is huge.”

— Hugging Face spokesperson

Amazon

large-scale AI model hardware requirements

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unverified Aspects and Performance Uncertainties

Independent benchmark results, safety assessments, and detailed licensing terms are not yet available. The actual performance of Inkling on practical workloads, especially in video processing and multimodal reasoning, remains unverified. It is also unclear how well the model handles real-world tasks or how effective the speculative prediction layers are in production environments.

Amazon

multimodal AI processing hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Testing and Evaluation

Developers and organizations will begin testing Inkling through supported inference engines, focusing on latency, accuracy, and hardware requirements. Expect upcoming independent evaluations, benchmark disclosures, and safety assessments to clarify its capabilities and limitations. Further, licensing details and potential fine-tuning applications are likely to be announced in the near future.

Building MCP Servers for AI Agents: Scalable Architecture Patterns, Security Design, and Production-Ready AI Infrastructure for Large Language Models

Building MCP Servers for AI Agents: Scalable Architecture Patterns, Security Design, and Production-Ready AI Infrastructure for Large Language Models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Inkling?

Inkling is a large, open multimodal AI model from Thinking Machines, capable of processing text, images, and audio, with 975 billion parameters.

Can Inkling process videos?

While the architecture supports image inputs with a temporal dimension, native video processing performance has not been evaluated, and no definitive claims about video capabilities are available yet.

Is Inkling suitable for running on personal computers?

Likely not. The model requires approximately 2 TB of VRAM for full deployment, making it impractical for typical consumer hardware. Hosted inference or specialized infrastructure is recommended.

What are the licensing and safety considerations?

Details about licensing, usage restrictions, safety evaluations, and training data are not yet disclosed, leaving questions about its commercial and research deployment open.

When will more performance data be available?

Further evaluations, benchmarks, and safety assessments are expected as developers and third parties test Inkling in real-world scenarios over the coming months.

Source: ThorstenMeyerAI.com

LABOR DAY SALES

Labor Day sales Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

10 Best Ultrawide Monitors for Work and Gaming in 2026

Discover the best ultrawide monitors in 2026 for work and gaming, including Dell, Samsung, MSI, and more, based on latest reviews and features.

The Controversy Surrounding Anthropic And The Alleged Theft Of Tens Of Thousands Of Songs

Anthropic faces a lawsuit from music publishers claiming unauthorized use of copyrighted lyrics in training its AI model, raising legal questions about fair use.

Getting 25 Gbps Thunderbolt Ethernet On My Mac Studio

A user reports successfully connecting a 25 Gbps Ethernet on Mac Studio using Thunderbolt technology, marking a significant upgrade in network performance.

Signal: Memory-Squeeze Check-In — Prices Are Cooling Because You’re Broke, Not Because It’s Fixed

Recent data shows memory prices are cooling because buyers are out of money, not because supply has increased, signaling a prolonged market squeeze.