At Capgemini Engineering, the world leader in engineering services, we bring together a global team of engineers, scientists, and architects to help the world’s most innovative companies unleash their potential. From autonomous cars to life-saving robots, our digital and software technology experts think outside the box as they provide unique R&D and engineering services across all industries. Join us for a career full of opportunities. Where you can make a difference. Where no two days are the same.
YOUR ROLE
We are looking for a Mid-Senior Data Scientist with a strong background in biomedical sciences and data engineering to join our growing team in Portugal. In this role, you will work at the heart of cutting-edge life sciences projects, helping to design, build, and govern the data foundations that power scientific discovery and product development. You will act as a bridge between scientific, engineering, and product stakeholders - translating complex biological knowledge into robust, scalable, and FAIR data solutions. In this role you will play a key role in:
- Designing and governing biomedical data models: leading source-to-canonical mapping, ontology alignment, and schema governance including versioning, changelogs, and downstream impact assessments to ensure data integrity and scientific accuracy across complex biomedical domains
- Building and maintaining data pipelines: developing robust, schema-driven pipelines in Python and SQL, performing exploratory data analysis, and implementing validation frameworks that support high-quality, reproducible scientific workflows
- Driving knowledge graph development: applying hands-on experience with RDF, OWL, SPARQL, and property graph modelling tools such as Neo4j and GraphDB to build and enrich knowledge graphs that connect biomedical entities across diverse data sources
- Applying machine learning and Gen AI: leveraging applied ML experience and familiarity with Gen AI tools (including text generation APIs, chatbots, and enterprise search solutions) to extract insight and value from scientific data at scale
- Championing FAIR data principles: designing and delivering FAIR data products, leading harmonisation efforts across multiple source systems, and ensuring persistent identifiers and provenance are embedded into every data product
- Aligning stakeholders across disciplines: driving alignment between scientific, engineering, and product teams through clear communication, structured documentation, and a solutions-focused mindset that keeps complex projects moving forward
- Working with biomedical ontologies and controlled vocabularies: applying deep knowledge of resources such as Ensembl, UniProt, and Gene Ontology, including judgment on when and how to extend or map them to real-world data challenges