Definition of an Inference Server
An inference server is software that enables artificial intelligence (AI) applications to communicate with large language models (LLMs) to generate data-driven responses. This process is called inference and represents the moment when the final result is delivered—where the business actually captures value.
To operate efficiently, LLMs require substantial storage, memory, and infrastructure resources to perform inference at scale, which explains their potentially high cost.
Success in AI strategies depends on the hardware and software supporting inference capabilities. Red Hat AI Inference Server optimizes inference to allow your teams to scale while maintaining cost-effectiveness.
How Red Hat AI Inference Server Works?
Red Hat AI Inference Server enables fast, cost-effective inference at scale. Being open source, it supports all generative AI models, all AI accelerators, and any cloud environment.
Based on vLLM, this inference server optimizes GPU usage and shortens response times. Combined with the LLM Compressor tool, it improves inference efficiency without compromising performance. With its cross-platform adaptability and a growing contributor community, vLLM is becoming the Linux® of generative AI inference.
Models of Your Choice
Red Hat AI Inference Server supports all major open-source models with flexible GPU portability. You can even run models beyond text and code, such as geospatial data models that can interpret your physical environment.
You can use any type of generative AI model or choose from our curated collection of third-party open-source models, optimized and validated for efficient execution on the Red Hat AI platform.
Model validation in Red Hat AI is performed using open-source tools such as the GuideLLM framework, Language Model Evaluation Harness, and vLLM to ensure reproducibility.
Our expert consultants are ready to listen, understand your needs, and deliver the perfect solutions for your business.