Data Engineer for AI Stack (f/m/d) @ A1 Competence Delivery Center
А1 България ЕАД Top работодател
над 300 служителя
Data Engineer for AI Stack (f/m/d) @ A1 Competence Delivery Center
София
длъжност на пълно работно време

Data Engineer for AI Stack (f/m/d) @ A1 Competence Delivery Center

София длъжност на пълно работно време

Описание на позицията

Strength. Care. Growth

A1 Competence Delivery Center is a vital component of A1’s telecommunications business. Acting as an expertise hub, CDC is dedicated to delivering a full range of high-quality IT, network, financial and other services to support A1’s operations across all OpCos, independent of location.

Using the power of being OneGroup and leveraging synergies, CDC enables transparency of resources, key skills and knowledge expansion and personal career growth opportunities’ enhancement, paired with job stability.

This job can be performed by all countries within our A1 footprint.
Job Purpose:

As a Data Engineer for AI Stack, you will make enterprise data usable by the AIaaS platform. Your focus will be data ingestion, transformation, orchestration, metadata, lineage, replay, quality, and secure integration with A1 data sources.

You will work across APIs, databases, files, object storage, event sources, and analytical platforms. You will help establish reusable integration patterns that can support AI, machine-learning, RAG, analytics, and operational automation use cases.

Role insights:

  • Design and implement batch, incremental, API-based, and file-based data-ingestion pipelines.
  • Integrate the AIaaS platform with enterprise systems, databases, object storage, SaaS services, data lakes, and approved on-premises sources.
  • Define processing patterns for transformation, enrichment, validation, deduplication, checkpointing, retries, replay, and backfill.
  • Implement idempotent pipelines with clear failure and recovery semantics.
  • Develop reusable connectors, Python jobs, orchestration workflows, and data-service APIs.
  • Establish source-to-target mappings, schemas, contracts, and validation rules.
  • Contribute to the design of operational storage, object storage, analytical layers, open-table formats, vector stores, and feature stores.
  • Integrate pipelines with orchestration, metadata, lineage, observability, and data-quality services.
  • Define how lineage is generated and propagated rather than relying solely on passive metadata cataloguing.
  • Support RAG ingestion, document processing, embedding pipelines, feature engineering, model-training data preparation, and batch inference.
  • Work with security and privacy stakeholders on data classification, minimisation, pseudonymisation, retention, and access controls.
  • Provide realistic workload profiles for capacity, storage, network, and performance planning.
  • Produce data-flow diagrams, interface catalogues, operational runbooks, and support documentation.
  • Help distinguish reusable platform capabilities from use-case-specific engineering.

What makes you unique:

  • Strong data-engineering experience using Python, SQL, and modern orchestration or integration tools.
  • Experience building reliable production pipelines across heterogeneous data sources.
  • Knowledge of ETL/ELT, APIs, relational databases, object storage, file ingestion, and schema evolution.
  • Understanding of data-quality controls, lineage, metadata, replay, backfill, and idempotency.
  • Experience with workflow orchestration and automated deployment.
  • Familiarity with containerised execution environments and Kubernetes-based data workloads.
  • Ability to analyse failure modes and design recoverable data processes.
  • Strong communication skills and the ability to work with source-system owners, architects, platform engineers, governance teams, and use-case developers.

Nice to have:

  • Airflow, Airbyte, Meltano, OpenMetadata, dbt, Spark, Trino, Dremio, DuckDB, Iceberg, or Delta Lake.
  • MLflow, Feast, vector databases, embedding pipelines, or RAG ingestion.
  • Azure, Exoscale SOS/S3, Cloudera, Synapse, Teradata, PostgreSQL, Oracle, or Microsoft SQL Server.
  • Streaming technologies such as Kafka or Flink.
  • Experience with OpenLineage and OpenTelemetry.
  • Knowledge of AI/ML data lifecycles, feature engineering, or model-serving integration.

If you have any questions, please do not hesitate to contact:Mariya Ivanova

Job code: AEM060P210
Job classification:P2-10,KV4, PT 3/1