IBM and artificial intelligence startup Together AI have signed a $240 million multi-year agreement to build a large-scale AI inference cluster on IBM Cloud. The project, announced by IBM on August 11, will use Nvidia HGX B300 systems and Spectrum-X Ethernet networking. The companies expect the cluster to become available in the first quarter of 2027.
The new infrastructure is designed to run inference for open-source AI models. Inference is the stage in which a trained model is used to generate responses to new inputs. As companies move more AI applications into production, the computing capacity required for this stage has become a major focus for cloud providers and chipmakers.
HGX B300 and Spectrum-X will work together
IBM’s official announcement says the cluster will combine Nvidia HGX B300 systems with Spectrum-X Ethernet networking. The companies describe it as the first dedicated large-scale inference cluster on IBM Cloud built with HGX B300 systems. The infrastructure is intended to support Together AI’s delivery of open-source model services to enterprise customers. The announcement did not provide a final capacity figure or a detailed deployment plan.
The agreement’s value and planned availability window were confirmed by IBM. However, the company also included its standard caution that forward-looking goals and intentions may change. The first quarter of 2027 should therefore be treated as the current target rather than a guaranteed launch date.
HGX B300 is a server platform that combines multiple accelerators with high-speed interconnects. Spectrum-X is the Ethernet networking technology that carries data between systems. Using the two components together supports the goal of scaling inference workloads in production. IBM says the systems are designed for performance and efficiency, while real-world capacity and utilization can only be measured after the deployment enters service.
Open-model enterprise workloads are the focus
Together AI plans to use the capacity to provide inference services for open models. Its platform allows companies to train models and run inference workloads using open technology. Reuters reported that open models have attracted businesses seeking to manage costs and reduce dependence on closed systems.
According to IBM’s release, Together AI’s platform covers inference, training, fine-tuning and agent-based workflows. Together AI says its inference product now serves 400 trillion tokens per month. Because this figure comes from the company, it should be understood as a reported internal metric rather than an independently measured total.
A strategic capacity move for IBM Cloud
The project illustrates that competition in AI infrastructure extends beyond developing models. Production inference also requires accelerators, high-speed networking and reliable cloud operations. IBM will combine Nvidia hardware with its enterprise cloud services, while Together AI will use that capacity to deliver open-model services to customers.
Together AI recently announced an $800 million Series C financing round at an $8.3 billion valuation, a figure also reported by Reuters. The new $240 million agreement is not another funding round. It is a commercial arrangement covering infrastructure deployment and capacity on IBM Cloud.
The outcome will depend on whether the systems are deployed on schedule and on enterprise demand for open-model inference. For now, the verified elements are the size of the agreement, the hardware family, the location of the initial deployment in the United States and the target availability in the first quarter of 2027. Performance, utilization and customer impact can only be assessed after the cluster begins operating.
