About this role
About the Role
We’re looking for a Senior Data Engineer to join the Research Engineering Team and play a pivotal role in shaping the data infrastructure that powers Cyera’s innovative solutions. In this role, you’ll design, build, and scale data pipelines and data warehouses from the ground up, creating the backbone for the world’s leading data classification engine. You’ll collaborate closely with researchers and machine learning experts to enable the development of advanced models, ensuring a seamless flow of high-quality, reliable, and well-structured data. Your contributions will directly impact the scalability, performance, and precision of our platform, advancing our mission to protect critical data. What You'll Do
-
Design and build scalable batch and streaming data pipelines that process complex datasets from diverse sources, enabling reliable and high-performance model training and inference.
-
Collaborate closely with researchers and data scientists to deliver high-quality, structured datasets that accelerate experimentation and model iteration.
-
Lead large-scale historical backfills and migration initiatives to ensure data consistency and integrity across evolving storage and compute platforms.
-
Optimize data workflows through advanced query tuning, indexing, partitioning, and cost optimization strategies to support efficient large-scale analytics.
-
Architect and maintain high-performance cloud-based data platforms using modern data stack components across AWS, GCP, or Azure.
-
Operate distributed data processing engines to handle massive volumes of structured and unstructured data.
-
Design and implement event-driven architectures using large-scale queue systems such as Kafka or SQS to ensure reliable and efficient data movement.
-
Develop automated monitoring and validation systems that guarantee uptime, schema compatibility, and pipeline reliability.
-
Deploy and manage data infrastructure in containerized, Kubernetes-based environments to support scalable and resilient services. What We're Looking For
-
5+ years of experience in software engineering, with 2+ years focused on data engineering, building and operating large-scale data platforms.
-
Proven experience designing and optimizing data pipelines, data warehouses, and big data solutions.
-
Strong proficiency in data and advanced SQL, including: Complex analytical queries, Query performance tuning, Indexing & partitioning strategies
-
Experience working with Relational and NoSQL databases
-
Large-scale queue systems (Kafka, SQS, etc.)
-
Experience working with Distributed processing engines
-
Strong background in distributed systems and event-driven architectures.
-
Experience working with cloud-native infrastructure and high-scale systems.
-
Practical experience deploying and operating services in Kubernetes-based environments.
-
Ability to thrive in a fast-paced research environment, solving complex data challenges with scalable and innovative solutions. Nice to Have
-
Experience with additional data processing frameworks beyond the core stack.
-
Background in cybersecurity-related data infrastructure projects.
-
Familiarity with advanced machine learning workflows and their unique data challenges.
-
Contributions to open-source data engineering tools or frameworks.
-
Data lakes and modern storage formats (Delta Lake, Iceberg, Snowflake) Compensation & Benefits