๐ค๐ Building AI Behind Closed Doors: How Evrone Runs Private LLM Systems
Table of contents
๐ค๐ Building AI Behind Closed Doors: How Evrone Runs Private LLM Systems

โ๏ธ๐ข From Cloud Dependence to Full Control: Evroneโs On-Prem AI Blueprint
Table of contents
- ๐ About the Project
- ๐งฉ The Core Challenge
- ๐ฅ๏ธ Hardware Matters
- โ๏ธ Software Complexity
- ๐ง Model Testing & Speed
- ๐๏ธ From Setup to Production
- ๐ฅ The Right Team
- ๐ Final Thoughts
Modern companies want AI power without sending sensitive data outside their walls. One client asked Evrone to build a fully private assistant that worked entirely inside internal infrastructure. No public APIs. No external cloud dependency. Full ownership.
The assistant needed to answer natural-language requests, automate repetitive tasks, integrate with internal tools, and support agent workflows. Evrone designed the system so every prompt, file, and response stayed inside the clientโs environment.
๐งฉ The Core Challenge
Evrone focused on three essential tasks:
- Choosing hardware that could run serious workloads reliably.
- Building the software layer for orchestration, deployment, and updates.
- Testing models for speed, stability, and compatibility.
Private AI is never just โinstall a model and go.โ Evrone treated it as full infrastructure engineering.
๐ฅ๏ธ Hardware Matters
For this case, Evrone used an enterprise server with 8ร NVIDIA H100 GPUs. That setup handled sustained production traffic, not just demos.
Still, Evrone recognized that not every company needs that scale. Smaller workloads can run on compact machines such as Mac Studio systems or lighter GPU servers. Architecture should match the real business case.
Kubernetes gave the platform room to grow horizontally across multiple nodes.
โ๏ธ Software Complexity
Hardware sets limits, but software defines the experience. Evrone evaluated tools such as:
- vLLM
- Ollama
- llama.cpp
- mistral-rs
- SGLang
The open-source ecosystem remains fragmented. Some models prefer Safetensors, others use GGUF, while Apple devices often rely on MLX. Evrone selected combinations that stayed stable in Linux production environments.
๐ง Model Testing & Speed
Evrone tested several models, including GLM and Qwen. Qwen 3.5 became the best fit because it balanced quality and compatibility.
Performance changed everything:
- โ ๏ธ 20 tokens/sec felt slow in real workflows.
- โ 160 tokens/sec created smooth responses.
That speed made agent chains and multi-step reasoning practical.
๐๏ธ From Setup to Production
Evrone built the core system in 3 weeks, then used one more week for:
- Profiling performance
- Preparing models
- Validating agent integrations
Today the platform runs in production with GitOps, Kubernetes, and automated delivery pipelines. The client team now manages it independently.
๐ฅ The Right Team
Evrone found that private AI does not need a huge staff. It needs the right mix:
- Software engineers for integrations and workflows
- ML engineers for models and prompts
- DevOps engineers for reliability
- Business analysts for real use cases
๐ Final Thoughts
Private AI has matured. Evrone proved that on-prem LLM systems can power real operations, not just experiments. Companies trade some peak cloud speed for something more valuable: control, privacy, and long-term flexibility.
For finance, government, healthcare, and enterprise environments, that tradeoff often makes perfect sense.
๋ฉํ๋ฐ์ดํฐ
- post_id
- f74304d64ac9
- slug
- building-ai-behind-closed-doors-how-evrone-runs-private-llm-systems-f74304d64ac9
- url
- https://medium.com/evrone-en/building-ai-behind-closed-doors-how-evrone-runs-private-llm-systems-f74304d64ac9
- canonical_url
- https://medium.com/evrone-en/building-ai-behind-closed-doors-how-evrone-runs-private-llm-systems-f74304d64ac9
- author_url
- https://medium.com/@ekaterina_evrone
- status
- ok
- fetched_at
- 2026-06-10 15:53:41