Red Hat AI Inference 3.5

Early Access

Get started

Plan

Inference Operations

Deploy Distributed Inference with llm-d on Openshift Container Platform

Deploy and serve large language models at scale on Openshift Container Platform

Deploy Distributed Inference with llm-d on Azure or CoreWeave Kubernetes Service

Deploy Distributed Inference with llm-d on Azure or CoreWeave Kubernetes Service

Deploy the standalone Red Hat AI Inference container in OpenShift Container Platform

Deploy the standalone Red Hat AI Inference container in OpenShift Container Platform clusters that have supported AI accelerators installed

Deploy the standalone Red Hat AI Inference container in a disconnected environment

Deploy Red Hat AI Inference in a disconnected environment using OpenShift Container Platform and a disconnected mirror image registry

Monitor and troubleshoot Distributed Inference with llm-d deployments

Monitor and troubleshoot Distributed Inference with llm-d deployments

Inference serving language models in OCI-compliant model containers

Inferencing OCI-compliant models in Red Hat AI Inference

Speculative decoding

Speculative decoding with Red Hat AI Inference

Inference serving Mistral 3 models

Inference serving Mistral 3 models with Red Hat AI Inference

Inference serving geospatial foundation models

Inference serving geospatial foundation models with Red Hat AI Inference

Red Hat AI Model Optimization Toolkit

Compressing large language models with the LLM Compressor library

vLLM server arguments

Server arguments for running Red Hat AI Inference

Extending Red Hat AI Inference with tool calling capabilities

Configuring tool calling and chat templates for AI Inference

Additional Resources