Site Reliability Engineer (AIOps) (f/m/d) @ A1 Competence Delivery Center
А1 България ЕАД Top работодател
над 300 служителя
Site Reliability Engineer (AIOps) (f/m/d) @ A1 Competence Delivery Center
София
длъжност на пълно работно време

Site Reliability Engineer (AIOps) (f/m/d) @ A1 Competence Delivery Center

София длъжност на пълно работно време

Описание на позицията

Strength. Care. Growth
You will know we are the right place for you, if you are driven by:

  • Opportunities to learn and build your career.
  • Meaningful work in a stable and fast-paced company.
  • Diversity of people, projects, and platforms.
  • A supportive, fun, and inspiring place to work.

Job Overview:
We're looking for an Site Reliability Engineer to build and optimize our monitoring ecosystem during a major infrastructure transformation. You'll ensure end-to-end visibility across Azure and Exoscale, leveraging modern observability and AIOps practices to improve reliability, reduce alert fatigue, and proactively identify issues before they impact our services.

This job can be performed by all countries within our A1 footprint.
Role Insights:

  • Design and optimize multi-cloud observability pipelines for metrics, logs, and traces across Azure and Exoscale.
  • Implement AIOps solutions for anomaly detection, event correlation, and faster incident resolution.
  • Ensure seamless monitoring and visibility during cloud migration from Azure to Exoscale.
  • Integrate observability into GitHub Actions CI/CD pipelines to detect deployment and performance issues.
  • Develop automation, self-healing solutions, and runbooks together with DevOps engineers.

What Makes You Unique:

  • Experience in DevOps/SRE with knowledge of telemetry, observability, or data analytics.
  • Strong Kubernetes and observability stack expertise (Prometheus, Grafana, OpenTelemetry, ELK/PLG).
  • Hands-on experience with AIOps platforms (Datadog, Dynatrace, New Relic) or custom monitoring solutions in Python/R.
  • Understanding of cloud migrations and resilient, cloud-agnostic monitoring practices.
  • Fluent English and the ability to translate complex technical insights into actionable improvements.