SUMMARY:
Python/Spark/AI Developer
POSITION INFO:
Role purpose & context:
Our client is looking for a Python/Spark/AI Developer to replat form legacy T-SQL onto Spark/Delta Lake pipelines and build the APIs and AI features around them, writing type-safe, spec-first Python.
Key roles & responsibilities:
- Build Spark/PySpark pipelines on Delta Lake, replat forming T-SQL into Spark SQL.
- Write type-hinted, tested Python with pytest suites and CI lint/type-check gates.
- Build REST APIs (Fast API) for pipeline jobs, with auth, idempotency, and status semantics.
- Validate migrated pipelines against the legacy system with parity evidence.
- Work spec-first against design docs/ADRs, documenting changes as you go.
Must-have technical skills / experience:- Python (3.12) as a primary language — modern idiomatic Python: type-hinted code that passes strict static analysis (pyright/mypy), Pydantic models, ABC-based provider patterns, packaging with modern tooling (uv or Poetry), Click or similar CLI frameworks.
- Apache Spark / PySpark — production experience building data pipelines on OSS Spark (not only a managed vendor platform): Data Frame API and Spark SQL, partitioning and performance tuning, understanding of driver/executor architecture and Spark Connect.
- Delta Lake or an equivalent Lakehouse table format (Iceberg/Hudi) — MERGE INTO, schema evolution, time travel, idempotent write patterns.
- SQL — strong, dialect-portable — able to read legacy T-SQL (stored-procedure-era logic) and re-express its semantics faithfully in Spark SQL; comfortable reasoning about hashing, surrogate keys, and deduplication logic in set-based terms.
- Automated testing with pytest — fixtures, markers, tiered suites (unit/integration/e2e); test-first habits and comfort being held to parity/regression evidence.
- REST API development — Fast API or equivalent: request validation, auth (HMAC or similar signed-request schemes), idempotency, job-status semantics.
- Docker-based development — working daily against a Compose stack (Spark cluster, object storage, metastore, databases).
- Git + CI discipline — GitHub flow, PR-driven work with lint (ruff), type-check, and test gates on every change.
Preferred / nice-to-have technical skills:- LLM integration engineering — building provider-neutral AI features behind an abstraction: prompt construction, structured output validation, local/self-hosted inference (Ollama, vLLM, llama.cpp, LM Studio) as well as hosted APIs; evaluation and guard-railing of model output used in data workflows.
- Data governance / privacy engineering — PII tokenization and hashing schemes, k-anonymity concepts, re-identification risk, POPIA/GDPR-adjacent data handling; multi-tenant isolation awareness.
- Lakehouse platform components — Hive Metastore, Trino, Apache Ranger; how catalogs, views, and row/column policies compose into a governed query surface.
- Azure data estate familiarity — ADLS Gen2, Synapse/ADF concepts (the legacy being replaced), Azure Key Vault, Service Bus.
- Observability instrumentation — structlog/structured logging, Open Telemetry metrics and traces, Open Lineage.
- Migration/parity experience — replat forming pipelines with byte/cell-level output comparison against a legacy system.
- Certifications — Databricks Certified Developer for Apache Spark or equivalent Spark credential.
Seniority and experience:- – Intermediate-to-senior. 4+ years professional Python development, with 2+ years building Spark (or comparable distributed) data pipelines in production. Must work autonomously against written design/requirements docs and ADRs — the codebase is heavily spec-driven, and changes are expected to arrive with tests and documentation, not just code.
Required qualifications:- Bachelor’s degree in Computer Science, Engineering, or equivalent experience.
- Spark credential advantageous; proven record shipping production Python/Spark pipelines and working independently in a hybrid/remote team.
Why you’ll love working the company They believe in taking care of their team and creating an environment where you can thrive. As part of the company, you’ll enjoy:
- Flexible Working Arrangements: Whether you are a night owl or an early bird, they offer hybrid and remote options to suit your lifestyle
- Comprehensive Benefits: From a wellness program to home office reimbursements and continuous learning opportunities, they have got you covered.
- Team Culture:Fun team-building activities, regular socials, and a supportive, inclusive culture that values transparency, accountability, and work-life balance.
- Performance Incentives: Competitive salaries, ESOP, and recognition for your hard work.