Thinky's Inkling 975B-A41B: The Multimodal Open-Source Model Product Builders Need to Watch
TL;DR
- Thinky's Inkling 975B-A41B is the new leading American open-source multimodal model under Apache 2.0 license, offering genuine commercial freedom without the restrictions plaguing other "open" models.
- The model combines vision and language capabilities at a scale that makes it production-viable for startups and enterprises who can't afford proprietary API costs or need on-premise deployment.
- Apache 2.0 licensing is the game-changer here—not just model performance but legal clarity for product builders who need to ship without worrying about license violations or usage restrictions.
- For AI product managers, this represents a strategic inflection point: the gap between open and closed models is narrowing fast enough that build-vs-buy decisions need immediate recalibration.
The Open-Source Multimodal Moment We've Been Waiting For
We're at an interesting juncture in AI product development. For the past two years, if you wanted to build anything serious with vision-language capabilities, you essentially had three options: pay OpenAI's API fees, pay Google's API fees, or compromise significantly on quality with whatever open models existed. That calculation just changed.
Thinky's Inkling 975B-A41B—yes, that's the full model name, and yes, we're all going to have to get comfortable with increasingly cryptic model nomenclature—represents something I've been watching for closely: a genuinely capable multimodal model that you can actually use without license anxiety or API dependency.
The model emerged from Thinky, and according to the detailed breakdown from Latent Space, it's currently sitting at the top of the leaderboard for American open-source multimodal models. But the leaderboard position isn't the story. The story is what this means for the next wave of AI products.
What Makes This Model Different
Let's talk architecture and capabilities, because the details matter when you're making build decisions.
The "975B" designation refers to the model's parameter count—975 billion parameters operating in a mixture-of-experts (MoE) architecture. For context, that's massive, but the MoE approach means you're not activating all 975 billion parameters for every inference. The "A41B" indicates approximately 41 billion active parameters during inference, which is the number that actually determines your compute costs and latency.
This architectural choice is clever product thinking embedded in model design. You get the representational capacity of a near-trillion parameter model with the inference costs closer to a 40B model. For product builders, that's the difference between "theoretically impressive" and "actually deployable."
The multimodal capabilities span vision and language, which means you can feed it images alongside text prompts and get coherent, contextual responses. This isn't novel—GPT-4V and Gemini have been doing this. What's novel is doing it with Apache 2.0 licensing and weights you can download.
The Apache 2.0 Advantage: Why License Matters More Than Benchmarks
Here's where I need to inject some opinion based on years of navigating AI product decisions: the license is more important than the benchmark scores.
I've watched too many product teams get excited about an open model's capabilities, start building, then discover six months in that the license restricts commercial use above certain revenue thresholds, or prohibits specific use cases, or requires derivative works to maintain the same license. These aren't theoretical concerns—they're product killers.
Apache 2.0 is permissive in the ways that matter for commercial products:
- No revenue caps or usage restrictions
- No requirement to open-source your application code
- No ambiguity about commercial deployment
- Patent grant protection
When you're pitching investors or enterprise customers, "we built on an Apache 2.0 model" is a completely different conversation than "we built on a model with a custom license that restricts commercial use in certain contexts." The former is a technical decision. The latter is a legal liability.
Thinky's choice to release under Apache 2.0 isn't just generous—it's strategically smart for ecosystem development. They're lowering the activation energy for builders to actually ship products, which creates a feedback loop of real-world usage, bug reports, and community contributions.
Performance Reality Check: Where It Excels and Where It Doesn't
Let's be honest about capabilities, because overpromising model performance is how product teams end up with disappointed users.
Based on the benchmarks and early testing reports, Inkling 975B-A41B performs competitively with models like GPT-4V and Claude 3.5 Sonnet on vision-language tasks—particularly on tasks involving visual reasoning, OCR from images, and multimodal question answering. That's legitimately impressive for an open model.
But "competitive with" doesn't mean "identical to." Proprietary models still have edges in certain areas:
- Instruction following consistency across edge cases
- Refusal behavior and safety guardrails (which, depending on your use case, might be a feature or a bug)
- Latency on hosted API endpoints (though this gap closes if you're comparing to self-hosted Inkling)
For product builders, the question isn't "is this model perfect?" but "is this model good enough for my specific use case?" And increasingly, for a growing range of use cases, the answer is yes.
Strategic Implications for Product Teams
If you're building AI products, here's how I'd think about Inkling 975B-A41B strategically:
For Early-Stage Startups
This model fundamentally changes your cost structure. Instead of paying per-token API fees that scale linearly with usage (and can become your largest cost center), you can self-host and pay for compute infrastructure that scales more predictably. The break-even point depends on your volume, but for many products, it's lower than you'd think.
The multimodal capabilities also lower the barrier to building vision-language products. Previously, you needed either significant capital to pay for API access or significant ML expertise to fine-tune open models. Now you have a third option: download weights, deploy, iterate.
For Enterprises
The on-premise deployment story is compelling. If you're in healthcare, finance, or any regulated industry where data can't leave your infrastructure, proprietary APIs are often non-starters. Apache 2.0 licensed models you can run internally change the calculus entirely.
Additionally, the ability to fine-tune without license restrictions means you can create specialized versions for domain-specific tasks. That's not just a nice-to-have—it's often the difference between an AI feature that's a demo and one that's production-ready.
For AI Product Managers
This is a forcing function to revisit your model strategy. If your product roadmap assumes continued dependence on proprietary APIs because open models "aren't good enough," that assumption needs stress testing. The gap is closing faster than most roadmaps account for.
I think we're entering a period where the default choice for new AI products should be open models, with proprietary APIs used strategically for specific capabilities where the gap is still meaningful. That's a reversal from even 12 months ago.
The Deployment Reality: It's Not Just About Model Quality
Here's where theory meets practice. Having access to model weights doesn't automatically mean you can deploy them effectively.
A 975B parameter MoE model, even with only 41B active, requires significant infrastructure:
- Multi-GPU setups for reasonable inference latency
- Careful quantization decisions to balance quality and resource usage
- Proper serving infrastructure (vLLM, TensorRT-LLM, or similar)
- Monitoring and observability tooling
This isn't insurmountable—companies like Together AI, Replicate, and Modal have built platforms that abstract much of this complexity. But it's also not "download weights, run on a laptop."
For product teams, this means your deployment decision tree looks something like:
- Use a hosted API service that offers Inkling (lowest ops overhead)
- Deploy yourself on cloud GPU instances (more control, more complexity)
- Run on-premise if you have the infrastructure team and regulatory requirements
Each path has different cost, latency, and complexity tradeoffs. The key insight is that you now have these options at all for a model at this capability level.
What This Signals About the Open Model Ecosystem
Zoom out for a moment. Thinky's Inkling isn't happening in isolation—it's part of a broader pattern.
We've seen Meta's Llama series push the boundaries of open language models. We've seen Mistral and others prove that small, well-funded teams can compete with major labs on model quality. We've seen the community rally around truly open licenses (Apache 2.0, MIT) versus restrictive "open" licenses.
Inkling represents the multimodal frontier of this trend. And if the trajectory holds, we should expect:
- More competitive open multimodal models in the coming months
- Increasing specialization (models optimized for specific modalities or tasks)
- Better tooling and infrastructure for deploying these models
- More products built on open foundations by default
For those of us building AI products, this is fundamentally good news. More options, more competition, more innovation. The proprietary model providers will need to compete on service, reliability, and cutting-edge capabilities rather than being the only game in town.
Practical Next Steps for Builders
If you're a product builder reading this and thinking "okay, what do I actually do with this information?", here's my recommendation:
Experiment now, even if you're not ready to switch. Download the model (or access it via a hosted service), run it against your existing use cases, compare outputs to what you're currently using. You need empirical data on whether it meets your quality bar.
Benchmark your costs. Calculate what you're currently paying for proprietary API access, then model what self-hosting would cost at your volume. The break-even point might surprise you.
Consider a hybrid approach. Use open models for high-volume, cost-sensitive tasks and proprietary models for cases where you need the absolute highest quality. This isn't all-or-nothing.
Watch the fine-tuning opportunity. If you have domain-specific data, the ability to fine-tune an Apache 2.0 model could be a significant competitive advantage. Start thinking about what that dataset would look like.
The Broader Competitive Landscape
It's worth contextualizing where Inkling sits relative to other options:
Versus proprietary APIs (GPT-4V, Claude 3.5 Sonnet, Gemini): Still behind on some benchmarks, but close enough that cost and deployment flexibility often outweigh the quality gap. The proprietary models maintain an edge on cutting-edge capabilities and consistent behavior.
Versus other open multimodal models: Currently leading among American open-source options with Apache 2.0 licensing. Models like LLaVA and others exist, but often with smaller parameter counts or more restrictive licenses.
Versus building your own: For most teams, using Inkling as a foundation is dramatically more practical than training a multimodal model from scratch. The compute costs alone for training at this scale are prohibitive for all but the largest organizations.
The competitive position that matters most is this: Inkling is good enough that it expands the viable solution space for product builders. That's the threshold that matters.
Looking Forward: What to Watch
The release of Inkling 975B-A41B is a milestone, but the story is still unfolding. Here's what I'm watching:
Community adoption and fine-tunes: The real test of an open model is whether the community builds on it. Watch for domain-specific fine-tunes, integration into popular frameworks, and products shipping with it as a foundation.
Performance improvements: Open models iterate fast. Expect updated versions, quantized variants optimized for different deployment scenarios, and community contributions that improve performance.
Competitive responses: How do other model providers (both open and closed) respond? Does this push proprietary providers to lower prices or improve capabilities? Do other labs rush to release competing open multimodal models?
Enterprise adoption: The real validation will be when large enterprises start building production systems on this foundation. That's a slower process than consumer products, but it's where the sustainable market is.
Final Thoughts: The Build-vs-Buy Calculation Has Changed
Here's my core take: if you're making AI product decisions in 2024 and beyond, you can no longer default to "we'll just use OpenAI's API" without seriously evaluating open alternatives.
Thinky's Inkling 975B-A41B isn't just another model release—it's a signal that the open ecosystem has reached a capability threshold where it's genuinely competitive for production use cases. The Apache 2.0 license removes the legal ambiguity that's held back adoption. The multimodal capabilities open up product possibilities that were previously gated behind expensive APIs.
This doesn't mean open models are the right choice for every use case or every team. But it does mean the decision tree has fundamentally changed. Cost structure, deployment flexibility, and customization capabilities now weigh much more heavily in the analysis.
For product builders, this is an opportunity. The teams that move quickly to understand and leverage these open models will have structural advantages—lower costs, more control, faster iteration cycles. The teams that ignore this shift will find themselves at a disadvantage as the ecosystem evolves.
The gap between open and closed models hasn't disappeared. But it's narrow enough now that it's not the dominant factor in most product decisions. And that's a fundamentally different world than the one we were in even a year ago.
If you're building AI products, the time to experiment with models like Inkling is now—not because you need to switch immediately, but because you need to understand what's possible so you can make informed strategic decisions as the landscape continues to evolve. The future of AI products is increasingly open. The question is whether you'll be ready for it.
Frequently Asked Questions
What makes Thinky's Inkling 975B-A41B different from other open-source multimodal models?
Inkling combines three key differentiators: Apache 2.0 licensing (providing true commercial freedom), competitive performance with proprietary models like GPT-4V on vision-language tasks, and a mixture-of-experts architecture that delivers near-trillion parameter capabilities with the inference costs of a 40B model. The Apache 2.0 license is particularly significant because it removes the commercial restrictions and ambiguity that plague many other 'open' models, making it genuinely viable for production deployment.
When does it make sense to use Inkling instead of proprietary APIs like GPT-4V or Claude?
Consider Inkling when you need on-premise deployment for regulatory reasons, when API costs are becoming a significant portion of your budget (especially at scale), or when you need to fine-tune for domain-specific tasks without license restrictions. The break-even point varies, but for many high-volume applications, self-hosting open models becomes cost-effective quickly. However, proprietary APIs may still be preferable for use cases requiring the absolute highest quality or most consistent instruction-following.
What infrastructure do I need to deploy Inkling 975B-A41B?
Deploying Inkling requires multi-GPU infrastructure due to its size, even with the mixture-of-experts architecture that keeps only 41B parameters active during inference. You have three main options: use a hosted service that offers Inkling (lowest operational overhead), deploy on cloud GPU instances with proper serving infrastructure like vLLM or TensorRT-LLM (more control but more complexity), or run on-premise if you have the infrastructure team and regulatory requirements. Most teams without existing ML infrastructure should start with hosted options.
How does the Apache 2.0 license impact commercial product development?
Apache 2.0 licensing provides critical advantages for commercial products: no revenue caps or usage restrictions, no requirement to open-source your application code, clear permission for commercial deployment, and patent grant protection. This eliminates the legal ambiguity that can kill products six months into development when teams discover license restrictions. For investors and enterprise customers, Apache 2.0 represents legal clarity rather than potential liability, making it a strategic advantage beyond just the technical capabilities of the model.