← Back to all jobs
Arcadia

Analytics Engineer, Life Sciences Delivery Operations

Arcadia

4h ago

0$175k - $200kDataUnited Stateshimalayas
Analytics-EngineeringData-EngineeringLife-Sciences-Data-EngineeringData-Pipeline-EngineeringHealthcare-AnalyticsAnalytics-EngineerSenior-Analytics-EngineerData-Analytics-EngineerSenior

Job Description

Why This Role Is Important to ArcadiaLife sciences customers depend on Arcadia's real-world data to power drug development, safety surveillance, and outcomes research. As LS deal volume accelerates, the engineering foundation underneath delivery, i.e. quality, automation, data transformation evolution, and scale must keep pace.This is a hybrid role at the intersection of data engineering, data analysis, and delivery operations. You'll refactor, scale, own, and operate an automated RWD data delivery pipeline via dbt/AWS architecture, serving as the primary technical point of contact for channel partners.You write production-grade PySpark and dbt one day and may facilitate a data inquiry the next. You care deeply about both the correctness of the code and the clarity of the answer it produces. You're as comfortable in a GitHub PR as you are in a partner meeting.This is a foundational engineering role in a growing LS organization. The right person will help build the team as the business scales.What Success Looks LikeIn 3 monthsDeep familiarity with the end-to-end LS pipeline-from ingestion through dbt transformation, de-identification, and delivery-including the current Snowflake-based scripts and what will replace themOwnership of the channel partner data inquiry queue; resolving standard requests independently by leveraging AI agents, closing out in writing and in accordance with SLAsFirst contribution to the delivery pipeline codebase: a new or refactored dbt model, a PySpark debugging fix, or a validated QC delivery configurationThorough understanding of the monthly delivery cycle: Argo orchestration, Snowflake execution, manifest generation, Datavant/HealthVerity/IQVIA tokenization, and delivery QCIn 6 monthsCore delivery endpoint configurations migrated from manual Snowflake runbook to config-as-code; existing channel partners delivered with minimal manual script executionContributing increasingly receptive metrics toward a data quality scorecard, tracking pipeline health, completeness, and refresh SLAs across all channel partnersPHI de-identification compliance implementation process owned end-to-end, with clear documentation of rules appliedStrong working partnerships established with platform engineering (Data Engineering, TechOps) with clear interfaces and shared standardsIn 12 monthsMonthly delivery cycle runs automatically; manual Snowflake execution eliminated; delivery cycle time reducedRecognized internally as the technical authority on the LS data engineering architecture and delivery pipelineTest suites, acceptance criteria, and release documentation authored for all major pipeline changesPotentially beginning to mentor a junior team member as the LS delivery organization growsWhat You'll Be DoingRWD DATA PIPELINE ENGINEERINGAuthor and maintain dbt models and PySpark transformation jobs, replacing ad-hoc Snowflake scripts with governed, version-controlled, tested codeDesign and implement delivery endpoint configurations as code-customer, delivery target (Snowflake, S3), cadence, cohort filters, incremental and full historical refresh methodsWrite production-grade Python and PySpark for data transformation, validation automation, and delivery pipeline components, including customer-specific data models and schema validation logicConfigure and maintain AWS S3 delivery paths, Apache Iceberg table structures, and file staging patterns for partner data deliveryPartner with platform engineering to build and extend Argo Workflows orchestration for automated delivery execution, eliminating the manual monthly Snowflake runbookImplement and maintain HIPAA de-identification compliance rules in pipeline code in accordance with ED certificates; coordinate certification updates when new data elements or rule changes affect certified productsDELIVERY OPERATIONS & DATA QUALITYCoordinate and execute monthly RWD deliveries across all active channel partners: delivery job execution, manifest generation and validation, tokenization workflows, and QCDefine and monitor delivery quality metrics: pipeline health, data completeness, referential integrity, refresh SLA tracking, and minimum volume thresholds; manage on-time delivery against a >=95% targetValidate dbt model outputs and pipeline changes against expected schema and counts; define acceptance criteria and execute UAT before changes reach production or channel partnersDATA INVESTIGATIONS & PARTNER SUPPORTOwn the channel partner data inquiry queue-triage, investigate, resolve, and communicate on data questions and discrepancies; you are the primary research contact for channel partnersConduct root cause analysis on anomalies in RWD, claims, and clinical feeds, distinguishing source-level issues from transform-layer failures, and communicate findings clearly in writingBuild and maintain repeatable query libraries, data dictionaries, and end-to-end pipeline documentation in Confluence, reducing one-off analytical effort and improving institutional knowledgeBuil