About this role
About the Role Join the Alexa Smart Home Decision Sciences team as a Data Engineer, building scalable analytics in a large data-warehouse environment and partnering with Product Management, Software, and Data Science to power data-driven decisions for Alexa-enabled devices. You will design ETL pipelines, data models, and dashboards while prioritizing data privacy and compliance. What You'll Do
- Work with the product and development teams within Alexa org to understand the product vision and requirements.
- Collaborate with Product Managers, BI engineers, Software Engineers and Data Scientists to design, implement and support high quality data products.
- Partner with cross functional teams across Devices organization to ingest relevant datasets into the Alexa Smart Home BI data-warehouse.
- Collaborate with other peer data engineers to build self-service data platforms that possess key capabilities like data discovery, data lineage, proactive data monitoring and security/compliance monitoring.
- Manage data infrastructure, including capacity planning, cost optimization, and performance tuning.
- Leverage and manage AWS services like Bedrock, Sagemaker, S3, Redshift, Athena, Kinesis, Lambda, Data Lake etc.
- Implement data pipelines using best practices in data modeling, ETL/ELT processes by leveraging AWS technologies and big data tools.
- Build data pipelines that support AI/ML use cases and enable integration with AWS AI services such as Amazon Bedrock and SageMaker to embed AI capabilities into production workflows.
- Collaborate with Data Scientists to adopt best practices in data system creation, data integrity, test design, analysis, validation, and documentation.
- Help continually improve ongoing data infrastructure processes, automating or simplifying self-service modeling and production support for stakeholders. What We're Looking For
- 3+ years of data engineering experience
- Experience with data modeling, warehousing and building ETL pipelines
- 4+ years of SQL experience
- Experience with AWS technologies like Redshift, S3, AWS Glue, EMR, Kinesis, FireHose, Lambda, and IAM roles and permissions
- Experience with non-relational databases / data stores (object storage, document or key-value stores, graph databases, column-family databases)
- Experience with big data technologies such as: Hadoop, Hive, Spark, EMR
- Experience with Apache Spark / Elastic Map Reduce
- Experience working with large language models (LLMs), including prompt engineering, model selection, and instructions fine-tuning to optimize model performance for analysis against large datasets
- Experience with scripting and API integration with AWS AI services such as Amazon Bedrock and SageMaker Nice to Have
- Data privacy and security/compliance monitoring experience
- Familiarity with data governance and data quality tooling
- Experience with cloud-based data privacy controls and auditability