Senior AI Reliability Engineer
EPAM Systems
EPAM's Operational Intelligence practice is developing a new capability called AI Reliability Engineering (AIRE), which applies SRE principles and cloud-native practices throughout the lifecycle of production AI/ML and LLM systems. As clients transition their GenAI applications and agentic systems from pilot to production, they find that traditional APM provides no visibility into token latency, cost per request, semantic drift, or hallucinations. This role exists to close that gap.
Your work will involve instrumenting, monitoring, and hardening production AI systems, establishing AI-native service level objectives, and creating accelerators and reference implementations that the practice can reuse across accounts. Note that this is an engineering position rather than an L1/L2 support role, and it does not require a 24/7 on-call rotation.
What You'll Get
Responsibilities
EPAM strives to provide its global team of over 62,350 professionals in more than 55 countries with opportunities for professional growth from day one of collaboration. Our colleagues are the source of EPAM's success, so we value cooperation, strive to always understand our clients' business and aim for the highest quality standards. No matter where you are, you will join a dedicated, diverse community that will help you realize your potential to the fullest.
Your work will involve instrumenting, monitoring, and hardening production AI systems, establishing AI-native service level objectives, and creating accelerators and reference implementations that the practice can reuse across accounts. Note that this is an engineering position rather than an L1/L2 support role, and it does not require a 24/7 on-call rotation.
What You'll Get
- A brand-new discipline within EPAM where you shape the approach rather than follow an existing playbook
- Focus on engineering work without 24/7 on-call responsibilities
- Sponsored certification and training programs (Anthropic/Claude, Databricks, AI & Data Observability learning paths)
- Exposure to multiple clients and a direct route into presales and solution engineering
Responsibilities
- Add AI telemetry to production LLM, RAG, and agentic applications using OpenTelemetry and APM-native AI monitoring tools
- Deploy distributed tracing across multi-model chains, agent workflows, and retrieval-augmented generation pipelines to identify systemic latency and failure points
- Establish and track AI-native SLIs and SLOs, including time to first token (TTFT), throughput, error and refusal rates, cost per request, semantic drift, hallucination boundaries, and contextual accuracy
- Implement structured semantic logging and prompt/response monitoring to support quality analysis
- Create and maintain evaluation loops for output quality and safety using golden sets, LLM-as-a-judge methods, and Ragas/DeepEval-style frameworks, integrating them into CI/CD and runtime environments
- Monitor token-based cloud spend, model API rate limits, and quota usage while driving AI cost optimization efforts
- Set up AI gateways to manage API load balancing, failover, and fallback models across multiple LLM providers
- Build guardrails covering prompt injection and jailbreak filtering, output compliance, and bias and safety constraints
- Develop detection, triage, restoration, and problem management workflows for AI incidents, incorporating autonomous AI agents into root cause analysis to parse logs, generate hypotheses, and correlate state changes
- Enable rollback, canary, and fail-safe patterns for model, prompt, and configuration releases, while maintaining reproducibility through versioning of data, code, prompts, and models
- Develop practice accelerators, reference architectures, and internal training materials, and contribute to presales activities and client assessments
- 4+ years of experience in SRE, DevOps, platform, or observability engineering, with hands-on exposure to production AI/ML or LLM workloads
- Strong grasp of SRE fundamentals, including Golden Signals, SLI/SLO definition, error budgets and burn rate, incident lifecycle, and ITIL basics
- Strong Python skills for building instrumentation, automation, and evaluation tools
- Hands-on experience with OpenTelemetry and at least one APM/observability platform such as New Relic, Datadog, Grafana LGTM stack, Splunk, or Elastic
- Production experience with at least one cloud platform (Azure preferred, AWS or GCP also acceptable) and Kubernetes
- Solid understanding of LLM application architecture, including prompts, embeddings and vector stores, RAG, and agent orchestration tools like LangChain/LangGraph or similar
- Awareness of MLOps concepts, including model lifecycle (training vs. inference), model endpoints, containerization, and deployment/rollback patterns
- Experience with Infrastructure as Code using Terraform, along with CI/CD tools such as Azure DevOps, GitLab CI, or GitHub Actions
- B2+ English proficiency, since the role involves direct client interaction and requires clear technical communication in both writing and speech
- Familiarity with AI-specific observability and evaluation tools such as Traceloop/OpenLLMetry, Langfuse, Arize Phoenix, Ragas, DeepEval, or MLflow
- Experience with distributed inference serving at scale, including vLLM, KServe, Ray Serve, Kubernetes-native LLM orchestration, and GPU capacity planning
- Knowledge of AI security practices, including the OWASP LLM Top 10, prompt injection defense, and guardrail frameworks like NeMo Guardrails or Llama Guard
- Experience with Databricks (including Mosaic AI/MLflow) or Azure AI Foundry
- Relevant certifications such as Anthropic Claude, Azure AI Engineer, AWS ML Specialty, or Databricks GenAI
- Background in FinOps for AI workloads, particularly token and GPU cost modeling
- Experience in Data Reliability Engineering, covering data quality and pipeline SLOs, since AI reliability depends on data reliability
- Prior mentoring or team lead experience
- With us you can:
- Work on a flexible schedule remotely or from any of our comfortable offices or coworking spaces in Ukraine
- Receive the necessary equipment to perform your work tasks
- Change projects and technology stacks within EPAM
- Gain experience in various business domains (Insurance, E-commerce, Healthcare, Finance, Travelling, Media, Artificial Intelligence, and more)
- Relocation opportunities may be available for eligible candidates, depending on the role and openings at other EPAM locations
- Participate in volunteer, charity programs and communities (both technical and interest-based)
- We focus on your professional growth:
- You can plan your individual career path together with your manager
- Receive regular feedback from colleagues
- Improve your English for free with certified teachers (Speaking Clubs, client interview preparation courses, etc.)
- Get the opportunity to undergo free training and certification in AWS, GCP, or Azure Clouds
- Use the internal E-learn training program (18,200+ specialized training and mentoring programs)
- Access corporate accounts on LinkedIn Learning, Get Abstract and other partner resources
- Study at EPAM Solution Architecture School with the instructors who are practicing architects
- Develop as a leader, join Delivery Management, Resource Management, Leadership Essentials school and more
- Participate in internal communities (500+ meetups, technical discussions, brainstorming sessions, online events and conferences annually)
- What we offer:
- Vacation and sick leave (including a sick leave without a medical certificate)
- A wide range of Voluntary Medical Insurance programs providing both medical treatment and various preventive options (including sports activities)
- Medical insurance for family members at corporate rates
- Company support during significant life events (childbirth or adoption, marriage, etc.)
- Support for psychological comfort: discounts on services from mental health specialists or coaches, thematic training
- E-kids program - a free programming language training program for EPAMers' children
EPAM strives to provide its global team of over 62,350 professionals in more than 55 countries with opportunities for professional growth from day one of collaboration. Our colleagues are the source of EPAM's success, so we value cooperation, strive to always understand our clients' business and aim for the highest quality standards. No matter where you are, you will join a dedicated, diverse community that will help you realize your potential to the fullest.
Як відгукнутися?
Щоб відгукнутися на цю вакансію, вам необхідно авторизуватися на нашому сайті. Якщо у вас ще немає облікового запису, будь ласка, зареєструйтесь.
Розмістити резюмеСхожі вакансії
Бариста-бармен у ресторан NOA (центр)
Tomatina, NOA, Emily, Poke Lulu, Una Pinsa, сім'я ресторанів,
Львів,
25 ₴
2 дні тому
Ми— це команда ресторанів з душею. NOA,Emily, Tomatina, Una Pinsa та Poke Lulu, ТАШ — шість унікальних концепцій, об'єднаних любов’ю до їжі, сервісу і людей. У кожному нашому закладі— свій характер, стиль і ритм. Але всіх нас єднає спільне: бажання створювати не просто страви, а враження.Мине просто готуємо їжу — ми створюємо досвід! Що ти отримаєш: Графік роботи — позмінний...
Senior System Engineer IRC302362
GlobalLogic,
Львів,
6 днів тому
Description The Digital Health organization is technology team which focused on next generation Digital Health capabilities which deliver on the Medicine mission and vision to deliver Insight Driven Care. This role will operate within the Digital Health Applications & Interoperability subgroup of the broader Digital Health team, focused on patient engagement, care coordination, AI, healthcare analytics & interoperability amongst other...
Packaging Data Specialist
Nestlé,
Львів,
2 тижні тому
UA, Lviv Hybrid work for Lviv region candidates . Nestlé Business Solutions (NBS) is a global team delivering smart, efficient solutions that keep Nestlé running worldwide. We combine technology and collaboration to simplify processes and create real business value. Are you ready to join a multinational company and a dynamic team? We’re excited to offer an opportunity for Packaging Data...