About this role
About the Role Fusemachines is seeking a Senior Data Engineer to architect, design, and implement scalable, high-performance data solutions. This remote, full-time role focuses on building end-to-end real-time and batch data pipelines and cloud-native architectures using AWS, GCP, Azure, Databricks, and Snowflake. What You'll Do
- Design, build, and optimize real-time and batch data pipelines end-to-end across ingestion, transformation, storage, and BI layers.
- Architect scalable data systems using cloud-native technologies (AWS, GCP, Azure, Databricks, Snowflake) and promote best practices in data modeling, performance tuning, and governance.
- Implement ETL/ELT pipelines with dbt, Airflow, Dagster, or native cloud orchestrators (Glue, Data Factory, Composer).
- Integrate data from diverse sources including APIs, RDBMS, NoSQL, flat files, and streaming platforms (Kafka, Kinesis, Pub/Sub).
- Model data using 3NF, Star, and Snowflake Schema and design Lakehouse/Warehouse architectures.
- Ensure data quality, security (RBAC, encryption), lineage, observability, and cost optimization while collaborating with product and analytics teams.
- Mentor and guide junior engineers; promote best practices in SDLC and DevOps. What We're Looking For
- 5+ years of hands-on data engineering experience in production environments.
- Proficiency in Python, SQL, PySpark/Apache Spark.
- Expert in data modeling and Lakehouse/Warehouse architectures.
- Experience building pipelines with dbt, Airflow, Dagster, or cloud orchestrators (Glue, Data Factory, Composer).
- Experience integrating data from APIs, RDBMS/NoSQL, flat files, and streaming platforms (Kafka, Kinesis, Pub/Sub).
- Deep expertise in at least one major cloud data ecosystem (AWS, Azure, GCP, Snowflake, or Databricks).
- Proficient in SDLC & DevOps: Git workflows, CI/CD pipelines, and IaC (Terraform/CloudFormation).
- Strong data governance focus: data quality, lineage, observability, and security (RBAC, encryption).
- Excellent communication and collaboration skills. Nice to Have
- Experience with Snowflake features (SnowSQL, Streams, Tasks, Snowpark) and cost optimization.
- Delta Lake, Unity Catalog, Delta Live Tables (DLT), and Spark optimization.
- Familiarity with cloud data services across AWS, GCP, and Azure (e.g., Redshift, S3, Lake Formation, Glue, Lambda; BigQuery, Dataflow, Dataproc, Pub/Sub; Synapse Analytics, Data Factory, Azure Databricks, Stream Analytics).
- Prior experience implementing enterprise data platforms with governance capabilities. Compensation & Benefits