About this role
About the Role Fusemachines is seeking a Senior Data Engineer to architect, design, and implement scalable, high-performance data solutions. You will own end-to-end data pipelines, both real-time and batch, and develop cloud-native architectures using AWS, GCP, Azure, Databricks, and Snowflake to enable enterprise AI and analytics. What You'll Do
- Design, build, and optimize end-to-end real-time and batch data pipelines.
- Architect scalable data architectures on cloud platforms; optimize ingestion, storage, processing, and BI.
- Collaborate with data scientists, product, and analytics teams to deliver reliable data products.
- Implement data governance, data quality, lineage, observability, and security (RBAC, encryption).
- Leverage dbt, Airflow, Dagster, or native cloud orchestrators (Glue, Data Factory, Composer) to orchestrate pipelines.
- Tune performance, monitor cost, and apply best practices for Lakehouse/Warehouse patterns (3NF, Star, Snowflake Schema).
- Ensure data reliability, security, and compliance across data ecosystems. What We're Looking For
- 5+ years of hands-on data engineering experience in a production environment.
- Strong proficiency in Python, SQL (complex queries, performance tuning), and PySpark/Apache Spark.
- Expert knowledge of data modeling (3NF, Star, Snowflake Schema) and Lakehouse/Warehouse architectures.
- Proven experience building pipelines using dbt, Airflow, Dagster, or native cloud orchestrators (Glue, Data Factory, Composer).
- Experienced in integrating data from diverse sources: APIs, RDBMS/NoSQL databases, flat files, and streaming platforms (Kafka, Kinesis, Pub/Sub).
- Deep expertise in at least one cloud data ecosystem (Snowflake, Databricks, GCP, Azure, or AWS) with practical knowledge of relevant tools (e.g., SnowSQL, Delta Lake, BigQuery, Synapse, Redshift).
- SDLC & DevOps: Git workflows, CI/CD pipelines (GitHub Actions, Azure DevOps, AWS CodePipeline), and IaC (Terraform/CloudFormation).
- Data governance: strong understanding of data quality, lineage, observability, and security (RBAC, encryption). Nice to Have
- Familiarity with the cloud data ecosystems listed in the role (Snowflake, Databricks, BigQuery, Synapse, Redshift) and related tooling such as Delta Lake, Unity Catalog, DLT, and Spark optimization.
- Experience working across enterprise domains and collaborating with cross-functional teams to deliver data products. Compensation & Benefits
- Not disclosed in the posting.