Senior Platform Engineer (f/m/d) @ A1 Competence Delivery Center
А1 България ЕАД Top работодател
над 300 служителя
Senior Platform Engineer (f/m/d) @ A1 Competence Delivery Center
София
длъжност на пълно работно време

Senior Platform Engineer (f/m/d) @ A1 Competence Delivery Center

София длъжност на пълно работно време

Описание на позицията

Strength. Care. Growth

A1 Competence Delivery Center is a vital component of A1’s telecommunications business. Acting as an expertise hub, CDC is dedicated to delivering a full range of high-quality IT, network, financial and other services to support A1’s operations across all OpCos, independent of location.

Using the power of being OneGroup and leveraging synergies, CDC enables transparency of resources, key skills and knowledge expansion and personal career growth opportunities’ enhancement, paired with job stability.

This job can be performed by all countries within our A1 footprint.
Job Purpose:

We are building an enterprise AI-as-a-Service platform that brings together reusable AI services, DataOps, MLOps, LLMOps, Kubernetes-based workloads, APIs, and secure enterprise integrations.

As a Senior Platform Engineer in Data DC, you will focus on the service and integration layer above the cloud foundation. You will help deploy, integrate, automate, and operationalise AI and data-platform services on Kubernetes.

Role insights:

  • Deploy and maintain containerised AIaaS, DataOps, MLOps, and LLMOps services on Kubernetes.
  • Integrate services such as workflow orchestration, data integration, metadata management, model tracking, feature stores, notebooks, model serving, RAG, and LLM observability.
  • Develop and maintain Helm charts, Kubernetes manifests, configuration templates, and deployment pipelines.
  • Establish repeatable deployment patterns across DEV, test, and production environments.
  • Configure namespaces, service accounts, RBAC, resource quotas, network policies, secrets, certificates, and application ingress.
  • Integrate platform services with enterprise APIs, databases, object storage, identity providers, monitoring systems, and internal data sources.
  • Implement application-level observability using metrics, logs, traces, health probes, dashboards, and actionable alerts.
  • Support vulnerability remediation, container-image governance, certificate renewal, secrets rotation, backup integration, and platform hardening.
  • Diagnose failures spanning Kubernetes workloads, service configuration, APIs, identity, networking, storage, and external dependencies.
  • Produce deployment documentation, technical runbooks, troubleshooting guides, and handover material.
  • Review vendor deliverables and identify hidden infrastructure assumptions, privileged dependencies, portability gaps, and operational risks.
  • Help translate pilot implementations into repeatable, supportable production services.

What makes you unique:

  • Strong practical experience operating applications on Kubernetes.
  • Sound knowledge of Helm, Kubernetes manifests, services, ingress, RBAC, secrets, ConfigMaps, persistent volumes, and network policies.
  • Experience with CI/CD and infrastructure or configuration automation.
  • Proficiency in at least one scripting or programming language, preferably Python, Go, or Bash.
  • Experience integrating APIs, databases, identity services, object storage, and enterprise applications.
  • Working knowledge of monitoring and observability concepts across metrics, logs, traces, dashboards, and alerting.
  • Ability to troubleshoot across application, Kubernetes, network, identity, and storage boundaries.
  • Clear technical communication and the ability to collaborate across engineering, security, architecture, and operations teams.

Nice to have:

  • Experience with Cilium, CRI-O, Traefik, cert-manager, Kyverno, Falco, Harbor, Velero, Prometheus, Grafana, Loki, or OpenTelemetry.
  • Exposure to Airflow, Airbyte, Meltano, OpenMetadata, MLflow, Feast, JupyterLab, LangFlow, vLLM, or RAG platforms.
  • Experience with GPU scheduling, NVIDIA GPU Operator, model serving, or AI inference workloads.
  • Familiarity with Exoscale or another European or sovereign-cloud provider.
  • Experience assessing workload portability between Kubernetes distributions.

If you have any questions, please do not hesitate to contact:Mariya Ivanova

Job code: AEM060P311
Job classification:P3-11, KV5, PT 2/3