Tool Guides

CoreWeave Forge Unifies AI Development Workflow

CoreWeave Forge connects training, inference, evaluation and agent development in one environment, letting teams run the entire AI improvement loop continuously.

Headline card: CoreWeave Forge Unifies AI Development Workflow
On this page
  1. What changed
  2. Why it matters
  3. What to test
  4. The conclusion

What changed

CoreWeave launched CoreWeave Forge on September 30, 2026, a development platform that runs the entire AI agent evaluation loop in one connected environment. According to the announcement, Forge unifies training, inference, evaluation and agent development while remaining open to any model, framework or cloud.

Forge connects five stages: run, observe, curate, improve, evaluate and repeat. The platform integrates Weights & Biases Models for experiment tracking, post-training expertise from OpenPipe, the open-source marimo notebook project, and CoreWeave’s own services. MasterClass and Canva are already building on the platform.

Several capabilities ship with the launch. CoreWeave ARIA, a coding agent that analyzes experiment and observability data, is now generally available. CoreWeave Agent Lens is a new observability tool for production agents that analyzes traces to surface insights and propose fixes. CoreWeave Sandboxes, which provide isolated environments for runs, are also now generally available. The platform includes CoreWeave Registry for versioning model checkpoints and agent configurations, CoreWeave Notebooks for development, CoreWeave Post-Training services including serverless reinforcement learning and model distillation, and CoreWeave Inference with a new RL Rollouts capability in preview.

Forge comes in three editions: Free, Pro and Enterprise. According to CoreWeave, teams can sign up for a 30-day free trial of the Pro tier and start without a procurement cycle.

Why it matters

Enterprises deploying AI agents face a continuous improvement challenge. As Susanne Seitinger, vice president of product marketing at CoreWeave, explained, agents must be run, observed, evaluated and improved continuously. The company notes that historically, model and agent development has been fragmented across separate vendor tools that were not built to work together, causing production signals to not feed back into training runs and experiments to not inform evaluations.

Consolidating these workflows into one CoreWeave Forge AI agent evaluation loop addresses a real pain point for machine learning teams. When development tools are disconnected, handoffs lose signal and cost time. Unifying the loop means training runs, experiment tracking, evaluations and agent traces live where the model or agent actually runs.

Chen Goldberg, executive vice president of product and engineering at CoreWeave, stated that as teams put models to work with their own data and workflows, gaps between model capability and system performance become clear. Engineering teams need to understand those gaps, identify signals that matter, improve the next version and measure results under real operating conditions as a single connected system.

Speed of learning is the primary metric that matters. According to Seitinger, the faster teams learn, the faster they ship and the faster they get value. The timing matters because the workload mix is shifting. Post-training and continuous improvement are becoming standard practice as enterprises deploy production AI workloads.

What to test

Before adopting CoreWeave Forge, teams should verify several vendor claims:

  • Agent Lens performance claims. CoreWeave claims Agent Lens detects 20 percent more critical failures and fixes issues at one-tenth the cost compared to a general-purpose frontier LLM. Teams should test this against their own agent workloads and benchmark the actual cost difference in their environment.

  • Serverless RL benchmarks. The company states serverless RL trains 1.4 times faster at 40 percent lower cost than self-managed setups. This benchmark should be validated against your specific training workloads, model sizes and infrastructure setup.

  • Integration openness. While CoreWeave claims Forge stays open to any model, framework and cloud, test whether switching between your preferred tools and CoreWeave’s integrated services requires engineering effort or adds friction to your workflow.

  • Data flow integrity. Verify that production traces and experiment results actually flow back into the next training run without manual steps or data loss. Test with a real agent evaluation cycle to confirm the loop closes.

  • Notebook portability. Test whether prototypes developed in CoreWeave Notebooks transition cleanly into training, evaluation and production without requiring rewrites or losing custom visualizations.

  • Registry versioning accuracy. Test CoreWeave Registry’s ability to track checkpoint lineage and agent configurations across many versions without conflicts or versioning errors.

  • Free tier boundaries. Understand what features or compute limits apply to the Free edition and what happens when you scale usage. Clarify which capabilities require Pro or Enterprise tiers.

The conclusion

CoreWeave Forge addresses a genuine problem: the fragmentation of AI development tools makes it hard to close the loop between production performance and model improvement. By unifying training, evaluation, inference and agent monitoring in one platform, the company removes friction that slows down iteration.

The timing matters. Enterprises are moving production workloads into AI agents, and the workload mix is shifting toward continuous post-training and improvement. CoreWeave’s emphasis on keeping the platform open across models and frameworks reduces lock-in risk compared to alternatives that require deeper vendor commitment.

What to watch: whether teams actually adopt Forge as their primary AI loop platform or whether they continue using tools from multiple vendors. The real test is whether the claimed cost and speed improvements hold up in production at scale. Also monitor whether CoreWeave’s partner network delivers recipes and playbooks that accelerate time to value, or whether teams still need to build custom integrations to make the platform work.

The three-tier pricing model removes upfront procurement barriers, which could accelerate adoption among teams that want to start small. Track whether that strategy drives meaningful customer growth or whether most teams quickly graduate to paid tiers. The openness to external models and frameworks is significant, but requires validation that this flexibility does not create integration complexity that undermines the simplification Forge promises.

AI Tool Herald may earn a commission from some links on this site. It never changes what we report or recommend. Affiliate disclosure