
OPEN POSITION
Data Engineer
Join the Lander Analytics team and be part of an innovative group focused on
transforming data into actionable insights. We’re seeking passionate individuals eager
to make an impact in data science and analytics.
Overview
Lander Analytics is a boutique data science and AI firm in New York. We've run the New York Open Statistical Programming Meetup for over a decade, host the New York and Government Data Science & AI Conferences, and write the books and courses people use to learn R and Python. We also build production systems for government agencies, finance, academic research, retail, manufacturing, insurance, energy, pharmaceuticals, and NFL front offices.
Scope
You'll own pipelines behind a federal data platform—hundreds of them, in production, moving financial data every day. We're looking for someone with a few years of data engineering under their belt who's ready to take responsibility for critical workloads, diagnose the failures that don't have obvious causes, and help shape how the platform evolves.
Key Responsibilities
Design, build, and maintain production ETL pipelines using Python, Spark, and AWS Glue/Airflow
Own pipelines end-to-end: development, testing, deployment, monitoring, and troubleshooting
Help migrate legacy pipelines (e.g., IBM DataStage, Oracle) onto the modern AWS stack
Contribute to and help evolve the team's shared Python library and ETL standards
Implement data quality checks and improve observability across the pipeline portfolio
Mentor junior engineers and review their work
Participate actively in Agile ceremonies and collaborate closely with the Technical Lead on architecture decisions
Communicate clearly with both technical and non-technical stakeholders
Requirements
3-5 years of experience in data engineering or a closely related field
Strong Python, PySpark and SQL skills
Hands-on experience with Apache Spark and a workflow orchestrator (Airflow/MWAA preferred)
Solid experience with AWS data services (Glue, S3, Athena, Redshift, Lambda, or similar)
Experience designing and maintaining production data pipelines at scale
Comfortable working in an Agile/Scrum environment
Strong communication skills and a track record of mentoring less experienced engineers
Ability to work closely with an entirely remote team
Must be a U.S. citizen and eligible to obtain Public Trust designation
Bonus Skills
Experience with Iceberg, Parquet, or other lakehouse table formats
Experience migrating legacy ETL platforms (IBM DataStage, Perl, Hadoop) to modern architectures
Experience with dbt
Working knowledge of Docker
Familiarity with data quality/observability tooling
Exposure to Amazon Bedrock or other AI/ML tooling in a data pipeline context
Prior experience on a GSA Schedule or other federal contract
Familiarity with FISMA or NIST 800-53 compliance requirements
Lander Analytics does not discriminate on the basis of race, creed, color, ethnicity, national origin, religion, sex, sexual orientation, gender expression, age, height, weight, physical or mental ability, veteran status, military obligations, or marital status.
In compliance with federal law, all persons hired will be required to verify identity and eligibility to work in the United States and to complete the required employment eligibility verification document form upon hire.