Together Link: Run Open Models Inside Claude Code
Together AI released Together Link, a free CLI to run open-source models like Kimi K3 inside Claude Code and other coding agents.

What changed
Together AI released Together Link, a free, MIT-licensed CLI tool now in beta. The tool connects six coding agents to open models hosted on Together AI’s infrastructure. Supported tools include Claude Code, Claude Desktop, Codex, ChatGPT Desktop, OpenCode, and Pi.
The installation process is straightforward. Users paste a single command on macOS or Linux, and the tool installs via Bun. A Together API key is required to authenticate. Once installed, users can launch their preferred agent directly from the terminal with commands like togetherlink claude or shortcuts such as tclaude.
The tool offers access to four open models: Kimi K3, GLM 5.3, GLM 5.3 Flash, and DeepSeek V4.1 Flash, each with 1 million token context windows. Kimi K3 costs $3.00 per million input tokens and $15.00 per million output tokens. GLM 5.3 costs $1.40 in and $4.40 out. Both DeepSeek V4.1 Flash and MiniMax M3 are priced at $0.30 in and $1.20 out.
Why it matters
Together Link addresses a practical problem for engineering organizations. Teams currently spend tens of thousands to millions of dollars monthly on closed-model APIs when every coding task, from simple one-line fixes to full rewrites, routes to the same premium model. The tool lets developers use affordable open models for routine work while preserving access to more capable options when needed.
The setup preserves existing workflows. No local proxy runs on users’ machines. Each tool talks directly to Together’s hosted gateway. Terminal agents receive temporary per-launch configuration that removes itself when the session ends. For Claude Desktop and ChatGPT Desktop, the tool creates a separate, reversible profile. Users can switch back with a single command.
The auto router is the default behavior and works by reading the first task in a session. Quick fixes go to low-cost models like GLM 5.3 Flash while harder problems route to more capable models. With an Anthropic API key, routing happens between Opus 5.5 and GLM 5.3. Without one, it routes between GLM 5.3 and GLM 5.3 Flash. Routing occurs only once per session, so prompt caching continues to work.
Inside Claude Code, the model menu maps agent model names to open alternatives. Opus runs Kimi K3, Fable runs GLM 5.3, Sonnet runs GLM 5.3 Flash, and Haiku runs DeepSeek V4.1 Flash.
Billing is transparent and session-specific. Each proxied session prints token and dollar totals on exit. Users can run togetherlink usage --last 7d to see gateway-tracked spending across the past week. In Claude Code, the status line shows estimated spend beside equivalent Opus costs. All billing flows through existing Together AI pay-as-you-go or credit pack accounts.
Together also notes it serves significant OpenRouter token volumes. As of September 30, 2026, the company served 40.8% of DeepSeek V4.1 Flash demand, 28.2% of GLM 5.3 Flash demand, and 23.1% of Kimi K3 demand, suggesting material proxy volume.
What to test
Before adopting Together Link in a production workflow, professionals should verify several claims and behaviors:
Model capability claims. Kimi K3 and GLM 5.3 target hard coding work, while GLM 5.3 Flash and DeepSeek V4.1 Flash cover everyday tasks. Test whether the router’s routing logic matches this division in your actual use cases. Verify that simple tasks stay on fast models and complex ones route correctly.
Cost savings. The company claims over 50% savings and 50 to 80% versus all-Opus 5.5 sessions. Calculate your current token spend on a premium model, then run a representative session through Together Link and compare. Track spend using the built-in togetherlink usage command.
Router behavior and consistency. Confirm that the auto router’s per-session routing decision is stable and predictable. Test whether routing works as documented with and without an Anthropic API key. Verify that prompt caching is preserved after routing decisions are made.
Integration reversibility. Test switching back to your original agent setup using the documented commands. Verify that your settings, history, and preferences remain intact after toggling Together Link on and off.
Platform coverage. Together Link is limited to macOS and Linux during beta. Test on your actual platform and confirm that the one-command install completes without errors.
Model list stability. Commands, routing, and the model list may change during beta. Periodically check Together’s pricing page to confirm model availability and rates have not shifted between your sessions.
The conclusion
Together Link solves a genuine cost problem by routing coding tasks to appropriately scoped models instead of sending everything to an expensive frontier API. The one-command setup and reversible profiles lower friction compared to manual configuration. The transparent per-session receipts and usage reporting let teams audit their actual spend.
The tool is deployable today on macOS and Linux, though its beta status means model availability and pricing may shift. Watch for three things: stable model availability and pricing as the tool exits beta, public benchmarks comparing the stated capability tiers (easy work on Flash, hard work on frontier models), and expansion to Windows during or after beta. The usefulness of the auto router depends entirely on its routing accuracy in your specific workflow, so test it early on a representative workload before rolling out widely.
AI Tool Herald may earn a commission from some links on this site. It never changes what we report or recommend. Affiliate disclosure
Related stories

CoreWeave Forge Unifies AI Development Workflow
CoreWeave Forge connects training, inference, evaluation and agent development in one environment, letting teams run the entire AI improvement loop continuously.

Liquid AI builds personal AI with device-level context awareness
Liquid AI's Liquid Context software enables on-device personal AI agents to improve after deployment within fixed hardware constraints, without cloud dependency.

Natura's $99 Interface Smart Ring Puts AI Agents on Your Finger
Natura launched Interface, a $99 smart ring combining AI agent control with health tracking, letting users summon tasks hands-free via voice and access agents 24/7.