Data Scientist

  • Closed
  • US Company | Large ( employees)
  • LATAM (100% remote)
  • 4+ years
  • Long-term · 40h/week
  • Enterprise software
  • Full Remote

Required skills

  • Python
  • SQL
  • Azure SQL Database
  • AWS
  • ETL
  • AI Technologies
  • LLM
  • Claude
  • ETL
  • Azure

Requirements

Must-haves

  • 4+ years of data science experience
  • Proficiency with Python to build data extraction and processing pipelines
  • Experience with SQL and cloud data warehousing (Azure SQL, AWS)
  • Experience integrating AI/LLM tools (e.g., Claude, ChatGPT) into extraction and document-processing workflows
  • Experience building applications that move and transform data across systems (e.g., ETL pipelines, document parsing, system-to-system validation)
  • Experience shipping production applications rather than purely analytical work
  • Ability to work independently and translate evolving, ambiguous requirements into working technical solutions
  • Experience in regulated environments (e.g., life sciences, clinical trials), particularly with study protocols and quality-controlled documentation
  • Experience with data validation and simulation testing to confirm consistency across systems
  • Experience designing database schemas for research or clinical study data
  • Strong communication skills in both spoken and written English

Nice-to-haves

  • Startup experience
  • Experience with BI tools (e.g., Power BI)
  • Experience partnering with quality or clinical operations teams to gather technical requirements from non-technical stakeholders
  • Bachelor's Degree in Computer Engineering, Computer Science, or equivalent

What you will work on

  • Design and build a pipeline to automatically extract relevant fields from source documents (Word, Excel, protocols), replacing manual transcription
  • Integrate AI/LLM-based extraction (e.g., Claude, GPT) to parse unstructured content and convert it into standardized electronic records
  • Build and maintain a centralized database to store curated, structured study data
  • Develop logic to generate finalized documents (Word, PDF) containing the components required for trial execution
  • Build cross-system validation logic to confirm extracted data matches downstream records, keyed against a test compendium
  • Structure and query curated datasets using SQL-based data warehousing (Azure SQL, AWS)
  • Partner with internal quality and data teams to align on data curation standards
  • Document technical approach and data schemas to support maintainability and future handoff