Data Scientist
- Closed
- US Company | Large ( employees)
- LATAM (100% remote)
- 4+ years
- Long-term · 40h/week
- Enterprise software
- Full Remote
Required skills
- Python
- SQL
- Azure SQL Database
- AWS
- ETL
- AI Technologies
- LLM
- Claude
- ETL
- Azure
Requirements
Must-haves
- 4+ years of data science experience
- Proficiency with Python to build data extraction and processing pipelines
- Experience with SQL and cloud data warehousing (Azure SQL, AWS)
- Experience integrating AI/LLM tools (e.g., Claude, ChatGPT) into extraction and document-processing workflows
- Experience building applications that move and transform data across systems (e.g., ETL pipelines, document parsing, system-to-system validation)
- Experience shipping production applications rather than purely analytical work
- Ability to work independently and translate evolving, ambiguous requirements into working technical solutions
- Experience in regulated environments (e.g., life sciences, clinical trials), particularly with study protocols and quality-controlled documentation
- Experience with data validation and simulation testing to confirm consistency across systems
- Experience designing database schemas for research or clinical study data
- Strong communication skills in both spoken and written English
Nice-to-haves
- Startup experience
- Experience with BI tools (e.g., Power BI)
- Experience partnering with quality or clinical operations teams to gather technical requirements from non-technical stakeholders
- Bachelor's Degree in Computer Engineering, Computer Science, or equivalent
What you will work on
- Design and build a pipeline to automatically extract relevant fields from source documents (Word, Excel, protocols), replacing manual transcription
- Integrate AI/LLM-based extraction (e.g., Claude, GPT) to parse unstructured content and convert it into standardized electronic records
- Build and maintain a centralized database to store curated, structured study data
- Develop logic to generate finalized documents (Word, PDF) containing the components required for trial execution
- Build cross-system validation logic to confirm extracted data matches downstream records, keyed against a test compendium
- Structure and query curated datasets using SQL-based data warehousing (Azure SQL, AWS)
- Partner with internal quality and data teams to align on data curation standards
- Document technical approach and data schemas to support maintainability and future handoff