AI Infrastructure Engineer
Qapitol is building the control layer for enterprise AI. That demands infrastructure that is fast, resilient and secure — not as an afterthought, but by design.
As AI Infrastructure Engineer you own the cloud backbone. You design, build and operate scalable infrastructure across AWS, GCP and Azure, and tune it specifically for AI-driven and SaaS workloads: GPU provisioning, inference-layer scaling, model-serving pipelines. You automate relentlessly and treat security as part of the build, not a layer bolted on afterward.
Who thrives here
You have 3–5 years in DevOps or cloud infrastructure and you are already comfortable across the major clouds. Kubernetes and Docker are daily tools, not line items on a CV. You care about root cause, not just uptime metrics. And you are genuinely curious about how AI is reshaping the infrastructure domain — AIOps observability, AI coding copilots, the cost and performance quirks of vector databases and embedding stores.
What you'll own
- Design, implement and manage scalable cloud infrastructure across AWS, GCP and Azure, including auto-scaling for high availability and performance.
- Set up, monitor and manage Kubernetes clusters; optimise Docker containers for scale and efficiency.
- Build, maintain and improve CI/CD pipelines; automate workflows to reduce manual effort and raise deployment velocity.
- Implement robust monitoring and troubleshoot critical issues; build proactive fixes to prevent downtime.
- Manage infrastructure as code and integrate security into DevOps processes.
- Deploy and scale infrastructure for AI/LLM and SaaS applications — GPU provisioning, inference-layer scaling, model-serving pipelines.
- Apply AIOps and AI-assisted observability tools for proactive incident detection.
- Partner with engineering teams to align DevOps practices with product goals and guide developers on deploy and scale best practices.
What it takes
- 3–5 years in DevOps, cloud infrastructure or related roles.
- Proven track record managing cloud services across AWS, GCP and Azure.
- Expertise in Kubernetes for cluster management and orchestration; strong hands-on Docker experience.
- Hands-on experience with CI/CD tools such as Jenkins, GitLab CI/CD or similar.
- Proficiency with IaC tools including Terraform and CloudFormation.
- Solid networking knowledge — load balancing, DNS, firewalls.
- Strong troubleshooting skills with a focus on root-cause analysis and prevention.
- Fluency with AI coding copilots (e.g. GitHub Copilot, Cursor, Claude Code) for authoring and reviewing IaC.
- Familiarity with cost and performance trade-offs for AI workloads — vector databases, embedding stores, model-caching layers.
Nice to have
- Prior experience at an AI or SaaS company, and familiarity with the challenges of deploying AI-driven or SaaS solutions.
We respond to every application within 2 business days.