Data Engineer (Unstructured Data & AI Pipelines) – Part‑time
Neurons Lab
Job description
About the role
Neurons Lab is looking for a part‑time Data Engineer/Data Scientist to build an unstructured‑first data platform for a European private investment group. The project spans eight to ten two‑week sprints, focusing on ingesting, normalising, and enriching communications, documents and archival data.
Key responsibilities
- Implement capture pipelines for calls, emails, Slack and other messengers, with speaker attribution and opt‑out controls.
- Back‑fill years of historical communications, parsing, deduplicating and dating each record.
- Develop document parsing for PDFs, scanned packs, spreadsheets and slide decks.
- Design and run entity‑resolution and record‑linkage across multiple identifiers.
- Build chunking, embedding and vector/graph store loading pipelines (pgvector, OpenSearch, Pinecone‑class).
- Implement incremental sync via connector layer, handling edits, deletions and rate limits.
- Attach access scope and provenance metadata at ingestion for permission‑aware retrieval.
- Run PII detection, redaction and retention logic, providing audit evidence.
- Orchestrate pipelines with Airflow or Step Functions, adding monitoring and alerting.
- Control cost and latency through batching, tiered storage and unit‑economics reporting.
Required profile
- 4+ years of data‑engineering experience, especially with unstructured or semi‑structured data.
- Proven ability to integrate multiple third‑party APIs and perform historical back‑fills.
- Experience building pipelines that feed LLM‑based retrieval systems.
- Comfortable handling sensitive personal data in regulated environments.
- Ability to work autonomously as the sole data engineer in a small, distributed pod.
Required skills
- Strong Python programming.
- Solid SQL knowledge.
- Airflow or equivalent orchestration tools.
- AWS and/or GCP data stack, including VPC deployments.
- Vector stores such as pgvector, OpenSearch, Pinecone‑class.
- Graph store integration.
- API integration with Google Workspace, Microsoft 365, Slack, CRM systems.
- Entity resolution / record linkage (deterministic and fuzzy).
- PII detection, redaction, encryption and retention practices.
- Understanding of GDPR and EU data‑residency requirements.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in Armenia.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
A question about this job?
Ask it here: you will get the full job summary by e-mail, right away.
Published 8 hours ago
Expires 1 month from now
9 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Neurons Lab