What's on the Data Engineer exam
Exam code: DEA-C01
Practice the real format
Do you think you are ready? Put your knowledge to the test with a free, timed practice exam that mirrors the Data Engineer format — with instant scoring, per-domain breakdowns, and full answer explanations.
Start a practice exam →The AWS Certified Data Engineer – Associate (DEA-C01) validates your ability to implement and manage data pipelines on AWS, and to monitor, troubleshoot, and optimize their cost and performance in line with best practices. The target candidate has roughly 2–3 years of data engineering experience plus 1–2 years of hands-on experience with AWS services.
Domains covered
The exam is organized into weighted domains. The percentages indicate roughly how much of the exam each domain represents, which is a useful guide for allocating study time.
Data Ingestion and Transformation 34%
Ingesting data from varied sources, transforming and processing it, orchestrating data pipelines, and applying programming concepts — using services such as AWS Glue, Amazon EMR, and Amazon Kinesis.
Data Store Management 26%
Choosing the right data store, designing data models, cataloguing schemas, and managing data lifecycles across services like Amazon S3, Redshift, and DynamoDB.
Data Operations and Support 22%
Operationalizing, maintaining, and monitoring data pipelines, analyzing data, and ensuring data quality with services such as Amazon Athena and CloudWatch.
Data Security and Governance 18%
Applying authentication, authorization, encryption, privacy, and governance, and enabling logging — using services like AWS Lake Formation, IAM, and Amazon Macie.
How to prepare
The exam is applied and service-heavy: expect scenarios that ask you to choose and combine the right AWS data services for ingestion, storage, and governance. Hands-on familiarity with Glue, Athena, S3, Redshift, and Lake Formation matters. As with other AWS exams, 50 of the 65 questions are scored and 15 are unscored trial questions. Timed practice papers help confirm your depth across all four domains.
Data engineering concepts mapped to AWS services
This exam is heavily scenario-based: a question will describe a data engineering need and ask you to select the right AWS service or combination of services. Use this reference table to lock in the concept-to-service mapping across all eight topic areas.
| Generic Concept | AWS Service | Purpose / What it does |
|---|---|---|
| Data Ingestion & Integration | ||
| Batch data ingestion | AWS GlueAWS DataSync | Moves and ingests batch data from on-premises and cloud sources |
| Streaming data ingestion | Kinesis Data StreamsAmazon MSK | Captures and processes real-time streaming data |
| Data transfer | AWS DataSyncAWS Transfer Family | Securely transfers files between on-premises and AWS |
| Database migration | AWS DMS | Migrates and continuously replicates databases with minimal downtime |
| Event-driven ingestion | Amazon EventBridgeAmazon SNSAmazon SQS | Captures events and builds event-driven data pipelines |
| API-based ingestion | Amazon API GatewayAWS Lambda | Collects data through REST APIs and serverless processing |
| Data Storage | ||
| Object storage / Data Lake | Amazon S3 | Durable, scalable storage for structured and unstructured data |
| Data warehouse | Amazon Redshift | Analytics data warehouse for large-scale SQL workloads |
| NoSQL storage | Amazon DynamoDB | Managed key-value and document database |
| Relational database | Amazon RDSAmazon Aurora | Managed relational databases for transactional workloads |
| Time-series database | Amazon Timestream | Stores and analyzes time-series data at scale |
| Data lake governance | AWS Lake Formation | Simplifies creation, security, and governance of data lakes |
| Data Processing & Transformation | ||
| Serverless ETL | AWS Glue | Creates, schedules, and monitors ETL jobs without managing infrastructure |
| Spark processing | AWS Glue Spark JobsAmazon EMR | Distributed big data processing using Apache Spark |
| Hadoop ecosystem | Amazon EMR | Managed Hadoop, Hive, Spark, Presto, and HBase clusters |
| SQL query on S3 | Amazon Athena | Serverless SQL queries directly on S3 data — pay per query |
| Stream processing | Amazon Managed Flink | Real-time analytics on streaming data using Apache Flink |
| Serverless compute | AWS Lambda | Lightweight event-driven data transformations |
| Workflow orchestration | AWS Step Functions | Coordinates multi-step data processing workflows |
| Data Catalog & Metadata | ||
| Metadata catalog | AWS Glue Data Catalog | Central metadata repository for all datasets |
| Schema discovery | AWS Glue Crawlers | Automatically discovers schemas and updates the Data Catalog |
| Data discovery | Amazon AthenaGlue Data Catalog | Searches and queries available datasets across the lake |
| Data Quality & Validation | ||
| Data quality checks | AWS Glue Data Quality | Validates completeness, accuracy, and consistency of data |
| Data profiling | AWS Glue Data Quality | Profiles datasets and identifies anomalies before processing |
| Data validation | AWS Glue ETL | Implements validation rules and schema checks during ETL |
| Analytics & Query Services | ||
| Interactive SQL analytics | Amazon Athena | Ad hoc querying of data stored in S3 — no cluster needed |
| Business Intelligence | Amazon QuickSight | Dashboards, reports, and visual analytics for business users |
| Data warehouse analytics | Amazon Redshift | High-performance structured analytics using SQL at scale |
| Search & log analytics | Amazon OpenSearch Service | Search, log analytics, and visualization via OpenSearch Dashboards |
| Data Security & Governance | ||
| Encryption key management | AWS KMS | Manages encryption keys for data at rest and in transit |
| Secrets management | AWS Secrets Manager | Securely stores and rotates database credentials and secrets |
| Identity & access management | AWS IAM | Controls authentication and authorization across all services |
| Fine-grained data permissions | AWS Lake Formation | Row, column, and table-level access control on data lake assets |
| Data auditing | AWS CloudTrail | Records API activity for governance, compliance, and forensics |
| Data classification | Amazon Macie | Discovers and protects sensitive and PII data stored in S3 |
| Monitoring & Operations | ||
| Pipeline monitoring | Amazon CloudWatch | Monitors ETL jobs, custom metrics, logs, alarms, and dashboards |
| Log management | CloudWatch Logs | Centralized log aggregation and filtering for AWS services |
| Operational events | Amazon EventBridge | Automates responses to pipeline events, failures, and schedules |
| Cost optimization | AWS Cost ExplorerAWS Trusted Advisor | Monitors, forecasts, and optimizes AWS spending across services |
| Automation & Orchestration | ||
| Workflow scheduling | AWS Glue Workflows | Orchestrates ETL jobs, crawlers, and conditional triggers |
| Serverless orchestration | AWS Step Functions | Coordinates complex multi-service data processing workflows |
| Event automation | Amazon EventBridge | Triggers data workflows based on events or cron schedules |
| Performance Optimization | ||
| Query optimization | Amazon RedshiftAmazon Athena | Partitioning, distribution keys, sort keys, compression, and workload management |
| Storage optimization | S3 Intelligent-TieringS3 Lifecycle Policies | Automatically moves data to lower-cost storage tiers over time |
| Caching | Amazon ElastiCache | Reduces latency for frequently accessed datasets and query results |
| Data Formats & Optimization | ||
| Columnar data formats | Apache ParquetApache ORC | Efficient columnar storage for analytical query performance |
| Data compression | GZIPSnappyZSTD | Reduces storage costs and improves query throughput |
| Data partitioning | S3 PartitioningGlue Partitions | Improves query performance by limiting the data scanned |
| Data Reliability & Recovery | ||
| Backup & recovery | AWS Backup | Centralized backup management and restore across AWS resources |
| Cross-region replication | Amazon S3 CRR | Replicates objects across AWS Regions for disaster recovery |
| High availability | Aurora Multi-AZRedshift RA3 | Ensures resilient, fault-tolerant data platforms with failover |
Exam formats and passing scores are updated periodically by AWS. Always confirm the current details on the official AWS certification page before booking your exam.
Ready to test yourself?
Do you think you are ready? Put your knowledge to the test with a free, timed practice exam that mirrors the Data Engineer format — with instant scoring, per-domain breakdowns, and full answer explanations.
Start a practice exam →