Today, Thinking Machines released Inkling, its newest open-weights model and first model family release. On the same day, Brainbase supports hosting Inkling for production agents.
That matters because model launches should not become infrastructure projects. When a model like Inkling appears, teams should be able to try it, route to it, evaluate it, monitor it, and deploy agents on it immediately.
That is exactly what Brainbase is built for.
What is Inkling?
Inkling is the newest model from Thinking Machines: an open-weights, Mixture-of-Experts model designed for text, image, and audio inputs. Thinking Machines describes Inkling as a broad, customizable foundation model with 975B total parameters, 41B active parameters, and a context window of up to 1M tokens.
In plain English: Inkling is a big, flexible model built for teams that want more control over how intelligence behaves inside real products and workflows.
It is also available to fine-tune on Tinker, Thinking Machines' training platform, which makes it especially interesting for teams building domain-specific agents.
How to try Inkling
If you are searching for Inkling, testing the Inkling model from Thinking Machines, or looking for a way to host Inkling in production, Brainbase is ready.
Start building at app.brainbaselabs.com.
Host Inkling in production
Brainbase now lets teams host Inkling-backed agents from day zero of the model release.
You can use Brainbase to deploy agents on Inkling, run them across real tools and workflows, and manage the operational layer around the model: routing, sandboxing, permissions, monitoring, evaluations, and scaling.
Brainbase also lets teams use Inkling across a wide range of agent harnesses. That matters because the harness is where the actual product behavior lives: prompts, tools, memory, permissions, evals, interfaces, and workflow logic. If every new model requires a new harness integration, testing gets slow and production gets messy.
With Brainbase, teams can try Inkling inside the same harnesses and workflows they already use, including OpenCode, Claude Code, and Codex, compare it against other models, and route production traffic when it is the right fit.
The model layer is moving fast. Brainbase makes the deployment layer move just as fast.
Inkling for AI agents
Inkling is especially interesting for agentic workloads because it is built around the things production agents actually need: tool use, multimodal input, controllable thinking effort, instruction following, and customization.
For teams building with AI agents, that opens up a useful set of options. Inkling can become another model in a production routing layer, a base for specialized internal agents, or a model to evaluate against existing frontier and open-weight options.
- Host Inkling agents for internal tools and workflows
- Use Inkling across a wide range of agent harnesses
- Route tasks to Inkling based on cost, latency, or quality
- Evaluate Inkling against other models in real production traces
- Monitor behavior, reliability, and usage over time
- Deploy Inkling across the same surfaces your teams already use
Inkling vs other agent models
The future of AI will not be one model. It will be many models, deployed across many agents, each selected for the job it is best at.
Inkling from Thinking Machines is a strong example of that future: open weights, multimodal inputs, customization, and a model roadmap built for people who want to make AI their own.
Brainbase is the cloud for that world. New model on day zero. Agents in production the same day.
