Data Engineer | Data Scientist – Unstructured Data & AI Pipelines
Neurons Lab · Chisinau
Job description
About the role
Neurons Lab is looking for a part‑time Data Engineer/Data Scientist to build the ingestion and context layer for a European private investment group. The role focuses on unstructured‑first data pipelines, entity resolution, and AI‑ready vector/graph stores.
Key responsibilities
- Stand up capture pipelines for calls, email, Slack and other messengers, with speaker attribution and opt‑out controls.
- Back‑fill years of historical email, Slack, board protocols, decks and portfolio updates, handling parsing, deduplication and correct dating.
- Develop document parsing for PDFs, scanned packs, spreadsheets and slide decks.
- Implement identity and entity resolution across multiple sources (Slack handles, mail aliases, calendar invites, portfolio company references).
- Build chunking, embedding pipelines and load vector + graph stores (pgvector, OpenSearch, Pinecone‑class) according to the architect’s ontology.
- Implement incremental sync via connector layer, avoiding full re‑crawls and handling edits/deletions.
- Attach access scope and provenance metadata at ingestion for permission‑aware retrieval and audit.
- Run PII detection, redaction and retention logic, providing evidence to the client’s security function.
- Orchestrate pipelines with Airflow or Step Functions, adding monitoring, alerting and cost‑control measures.
- Write runbooks for hand‑over to the client’s own team.
Required profile
- 4+ years of data‑engineering experience, including extensive work with unstructured or semi‑structured data.
- Proven ability to integrate multiple third‑party APIs and perform historical back‑fills.
- Experience building pipelines that feed LLM‑based retrieval systems.
- Hands‑on experience handling sensitive personal data in regulated environments.
- Comfortable working as the sole data engineer in a small, distributed pod with minimal supervision.
Required skills
- Strong Python programming and solid SQL.
- Unstructured‑data pipelines: parsing, normalisation, deduplication.
- Embedding/retrieval infrastructure: chunking strategies, vector stores (pgvector, OpenSearch, Pinecone‑class) and graph stores.
- API and connector integration at scale (Google Workspace, Microsoft 365, Slack, CRM systems).
- Entity resolution / record linkage (deterministic and fuzzy).
- Orchestration tools: Airflow, Step Functions or equivalents.
- AWS and/or GCP data stack, including private/VPC deployments.
- PII detection, redaction, encryption and retention practices.
- Clear written English for documentation and runbooks.
- Knowledge of GDPR, data lineage, provenance and audit patterns.
Questions fréquentes
Why are you reporting this job?
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
A question about this job?
Ask it here: you will get the full job summary by e-mail, right away.
Published 5 hours ago
Expires 1 month from now
5 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Neurons Lab
Chisinau
Related job offers
-
Expert principal responsabil de supraveghere TIC și securitatea informației
National Bank of Moldova Chisinau -
Data Analytics and Automation Specialist (Mid-Level)
Moldcell Chisinau -
Senior Technical Support Engineer
ZevetOne Chisinau -
Senior Java Engineer – Payment Services
Technology Oriented Kichinev -
Administrator de Sistem – Secția Reţea și Echipamente
OTP Bank, Moldova Kichinev