EuroHPC Grant Awarded: OPAIRS Trains Specialized AI Agents for Manufacturing
EuroHPC grant for OPAIRS: 5,000 GPU hours on Leonardo BOOSTER for industry-specific SLM fine-tuning, with the goal of cutting inference costs by 75% compared to cloud APIs.

Generic language models know nothing about maintenance strategies under DIN EN 13306. They know no bills of materials, no machine histories, and no company-internal abbreviations. What they deliver is a plausible average drawn from billions of training documents. For the shop floor, that is not enough. OPAIRS takes a different path: domain-specifically trained agents that know their field because they were trained on it. The grant from the EuroHPC Joint Undertaking now makes this step possible.
What EuroHPC Means and Why Leonardo

European infrastructure for European AI
The EuroHPC Joint Undertaking provides European companies and research institutions with compute time on supercomputers within the EU infrastructure. OPAIRS receives 5,000 GPU hours on Leonardo BOOSTER in Bologna, one of the most powerful computing facilities in Europe, with NVIDIA A100 GPUs and a peak performance of 174 petaflops. The grant is tied to specific technical and ethical requirements: traceability of training data, compliance with EU AI ethics principles, and human-in-the-loop governance. For OPAIRS, these requirements are not extra work but system architecture.
Three Specialized Agents, One Shared Goal
The fine-tuning is based on the Qwen3 model family (4B and 8B parameters) and the QLoRA method, a resource-efficient approach to supervised fine-tuning on industrial datasets. Three domain-specific agents are being trained:
- Maintenance and repair agent: Diagnosing error codes, deriving repair measures, and placing them within existing maintenance strategies. Training is based on machine manuals, DIN/EN/ISO standards, CMMS data, and structured maintenance logs.
- Project management agent: Task planning, status tracking, and prioritization based on company-specific workflow logic. The agent learns how projects in manufacturing really run, not how textbooks describe them.
- Document search and retrieval agent: Precision search in industrial knowledge bases using SQL and vector database queries. Optimized for semantic accuracy in technical documents.

From generic to specialized: What fine-tuning changes in practice
A base model answers questions with what it knows. A fine-tuned model answers questions with what applies in your own company. The difference lies not in model size but in the quality of domain adaptation. QLoRA makes it possible to train small language models with 4B to 8B parameters for domain-specific tasks without changing the model structure. After fine-tuning, the trained LoRA adapters are merged into the base model and then deployed on-premise directly via vLLM. No cloud endpoint, no external provider, no per-request API costs.
75% Lower Inference Costs Compared to Cloud APIs

Predictable costs instead of variable token pricing
A fine-tuned 4B model running on-premise on dedicated hardware has no variable per-request costs after the initial setup. No token pricing, no volume thresholds, no price changes from the provider. Capacity is fixed, and costs are predictable. The goal is concretely measurable: inference costs for standard industrial queries are to fall 75% below the level of leading cloud API providers.
- Scaling without cost growth: On-premise inference scales with the hardware, not with usage frequency. More requests do not mean a higher bill.
- Specialization beats size: A 4B model trained on maintenance outperforms a generic 70B model in this field, one that can do everything but nothing especially well.
- Customer knowledge as a training foundation: The trained agents are designed as a starting point for further fine-tuning. Customers can contribute their own machine data, maintenance histories, and process documents without any data leaving the company.
European Infrastructure, European Data Sovereignty
Using Leonardo BOOSTER does not contradict the OPAIRS on-premise philosophy. Training takes place on European infrastructure, the training data contains no personal or customer-sensitive content, and the finished model runs exclusively on-site at the customer. EuroHPC projects are subject to the EU data protection and ethics guidelines. For OPAIRS, this means that the entire lifecycle, from dataset through training to deployment, takes place under European frameworks and documented traceability.
Once fine-tuning is complete, the trained models will be integrated into the OPAIRS platform and form the inference layer for production systems from Q4 2026.
More insights

OPAIRS SQL Agent: Comparing Six LoRA Adapters for Industrial Databases
Six LoRA adapters, a 26-question catalog spanning PostgreSQL, T-SQL and Apache Iceberg databases, two external reference models: OPAIRS has systematically evaluated its SQL agent. Two Granite adapters lead the field. For production use, however, the deciding factors are not only answer quality but also speed, memory footprint, concurrency and the available context from ERP, MES, PLM and other industrial systems.
Read article
RTX PRO 4500 Blackwell: 3.4x LLM Throughput Through Runtime Optimization
Same GPU, same main model, up to 3.4x the output: OPAIRS Runtime 3 raises the throughput of GPT-OSS-20B on the RTX PRO 4500 Blackwell to up to 2,637 tokens/s. At the same time, the tests show why Qwen3.8-27B will not take over the production stack for now.
Read article