Lead Operational Intelligence Engineer
EPAM Systems
Our team is seeking a dynamic, highly experienced professional to fill the role of Lead Operational Intelligence Engineer.
This role involves taking charge of developing, maintaining, and enhancing our cloud-based Elastic & Observability Platform. The successful candidate will spearhead strategic initiatives, mentor a top-performing technical team, and maintain platform reliability while promoting innovation and self-service capabilities for platform users. On-call rotation duties for monitoring platform health and functionality are also part of this position.
Kindly note that this role supports remote work, but only from within Ukraine.
Responsibilities
EPAM strives to provide its global team of over 62,350 professionals in more than 55 countries with opportunities for professional growth from day one of collaboration. Our colleagues are the source of EPAM's success, so we value cooperation, strive to always understand our clients' business and aim for the highest quality standards. No matter where you are, you will join a dedicated, diverse community that will help you realize your potential to the fullest.
This role involves taking charge of developing, maintaining, and enhancing our cloud-based Elastic & Observability Platform. The successful candidate will spearhead strategic initiatives, mentor a top-performing technical team, and maintain platform reliability while promoting innovation and self-service capabilities for platform users. On-call rotation duties for monitoring platform health and functionality are also part of this position.
Kindly note that this role supports remote work, but only from within Ukraine.
Responsibilities
- Ensure observability and search platforms exceed business SLAs in terms of availability, functionality, performance, and security
- Deliver technical leadership when complex incidents arise and ensure prompt escalation of resolutions during on-call shifts
- Create and maintain thorough platform documentation, standard operating procedures, and knowledge-sharing materials
- Work with cross-functional teams, stakeholders, and vendors to manage operational needs, advance strategic initiatives, and handle installations, troubleshooting, and upgrades
- Drive improvements to platform features and self-service tools, including advanced Elastic Synthetics and automated chargeback processes
- Design and build proofs-of-concept to advance platform innovation, such as AI-driven observability, sophisticated data processing models, or migration to Kubernetes-based platforms
- Guide the construction, deployment, and upkeep of Elastic clusters using Infrastructure-as-Code tools such as Terraform and Ansible, and coach team members on best practices
- Manage platform lifecycle tasks, such as component upgrades, capacity planning, cost optimization, and adapting to new compliance requirements
- Regularly evaluate and optimize ELK stack performance, covering ingestion, indexing, and query tuning for large-scale environments
- Build and improve alerting and incident management processes by integrating advanced monitoring tools like Kibana Rules, Watchers, and PagerDuty
- Manage the ingestion, enrichment, backup, and restoration of large-scale platform data, optimizing data workflows along the way
- Direct and plan major operational events, including SSL certificate rotations, cluster migrations, and scalability optimization efforts
- At least 5 years of experience in Operational Intelligence, demonstrating leadership and technical skill in managing large-scale observability platforms
- Proven ability to design and oversee Elastic clusters within complex, multi-cloud environments
- Comprehensive knowledge of Elastic Stack components, including advanced setups of Elasticsearch, Kibana, and Logstash
- High-level skills in Infrastructure-as-Code tools such as Terraform and Ansible, with flexibility to work with tools like Jenkins CI or GitOps frameworks
- Strong Python scripting abilities for automating processes, handling data, and expanding platform interoperability
- Solid grasp of incident management frameworks and workflows using tools such as PagerDuty, Uptrends, and other enterprise monitoring platforms
- Demonstrated success in diagnosing and resolving intricate platform issues within strict SLA timeframes
- Strong skills in managing and scaling fault-tolerant platforms, ensuring performance, security, and compliance across large distributed systems
- Proven track record of mentoring team members, managing priorities, and serving as a liaison between technical and non-technical groups
- Strong English communication skills (B2+ level), both written and verbal, with an emphasis on technical communication
- Skills in Groovy scripting or advanced Linux administration experience to streamline platform operations
- History of enhancing observability workflows through custom integrations in tools such as Uptrends, PagerDuty, or Elastic
- Practical experience configuring advanced Elastic Synthetics for reliable monitoring and custom synthetic testing
- Background in leading strategic initiatives like AI-driven modernization, cloud-native migrations, or cost-saving observability improvements
- With us you can:
- Work on a flexible schedule remotely or from any of our comfortable offices or coworking spaces in Ukraine
- Receive the necessary equipment to perform your work tasks
- Change projects and technology stacks within EPAM
- Gain experience in various business domains (Insurance, E-commerce, Healthcare, Finance, Travelling, Media, Artificial Intelligence, and more)
- Relocation opportunities may be available for eligible candidates, depending on the role and openings at other EPAM locations
- Participate in volunteer, charity programs and communities (both technical and interest-based)
- We focus on your professional growth:
- You can plan your individual career path together with your manager
- Receive regular feedback from colleagues
- Improve your English for free with certified teachers (Speaking Clubs, client interview preparation courses, etc.)
- Get the opportunity to undergo free training and certification in AWS, GCP, or Azure Clouds
- Use the internal E-learn training program (18,200+ specialized training and mentoring programs)
- Access corporate accounts on LinkedIn Learning, Get Abstract and other partner resources
- Study at EPAM Solution Architecture School with the instructors who are practicing architects
- Develop as a leader, join Delivery Management, Resource Management, Leadership Essentials school and more
- Participate in internal communities (500+ meetups, technical discussions, brainstorming sessions, online events and conferences annually)
- What we offer:
- Vacation and sick leave (including a sick leave without a medical certificate)
- A wide range of Voluntary Medical Insurance programs providing both medical treatment and various preventive options (including sports activities)
- Medical insurance for family members at corporate rates
- Company support during significant life events (childbirth or adoption, marriage, etc.)
- Support for psychological comfort: discounts on services from mental health specialists or coaches, thematic training
- E-kids program - a free programming language training program for EPAMers' children
EPAM strives to provide its global team of over 62,350 professionals in more than 55 countries with opportunities for professional growth from day one of collaboration. Our colleagues are the source of EPAM's success, so we value cooperation, strive to always understand our clients' business and aim for the highest quality standards. No matter where you are, you will join a dedicated, diverse community that will help you realize your potential to the fullest.
Як відгукнутися?
Щоб відгукнутися на цю вакансію, вам необхідно авторизуватися на нашому сайті. Якщо у вас ще немає облікового запису, будь ласка, зареєструйтесь.
Розмістити резюмеСхожі вакансії
AI Engineer (Senior/Lead)
EPAM Systems,
Львів,
9 годин тому
We are looking for a Senior/Lead AI Engineer to design, build and operationalize production-grade AI agents, multi-agent workflows and AI/ML solutions using Python, Azure AI Foundry, Semantic Kernel / Microsoft Agent Framework, LangGraph, LangChain, AutoGen and Strands SDK. You will work alongside architects and data teams to take solutions from prototype to scalable, reliable production. Kindly note that this role...
Staff Engineer Systems (f/m/div)
Infineon Technologies,
Львів,
2 дні тому
#WeAreIn for jobs that impact everyone's life. What if your ideas could change the way the world connects, powers up, or thinks? As a Staff Systems Engineer in our Research & Development team, you'll have the opportunity to merge creativity with your technical expertise by shaping the future of technology, driving groundbreaking projects, and bringing new ideas to life. Are...
Lead QA Engineer IRC302484
GlobalLogic,
Львів,
3 дні тому
Description Our client builds diagnostic imaging equipment used in hospitals and clinics worldwide — mammography systems, biopsy devices, and related hardware that clinicians rely on every day. Behind that hardware sits a cloud platform that keeps tens of thousands of these devices connected: collecting operational data, powering remote diagnostics so support engineers can fix problems without an on-site visit, and...