Google Professional Data Engineer Exam Guide (2026)

If you can say whether the stem is a Beam job, a Spark cluster, or a SQL graph already sitting in BigQuery, whether the store is a warehouse, a lake file, or a wide-column key, and whether the scheduler is an Airflow DAG or a serverless workflow, you are reading the right outline.
Professional Data Engineer is the Google Cloud certification for people who collect, transform, store, and deliver data so other systems can decide from it. The Professional Data Engineer certification page is the source for length, fee, question count, validity, and recommended experience. The Professional Data Engineer exam guide is the source for the five scored sections and the product names used on the sitting.
The sitting is live. There is no prerequisite exam. Google recommends 3 or more years of industry experience, including 1 or more years designing and managing solutions on Google Cloud. This guide is for people scheduling the current standard exam, people who already hold the title and need the current task list, and people moving from Associate Cloud Engineer work into the data platform role.
Who this exam is for
Take it as a map of the Google Cloud data-engineer job the current exam guide scores. The audience profile asks for someone who can turn a business or regulatory need into a pipeline, a store, and an operations plan, then stay with that system through ingest, transformation, delivery, and cost control. The same profile expects fluency with current processing, cleaning, enrichment, and query-generation tools, and with the storage models those tools write into.
The responsibility list on the certification page is concrete. Design data processing systems. Ingest and process the data. Store the data. Prepare and use data for analysis. Maintain and automate data workloads. Those five verbs are the five scored sections. The sitting punishes people who treat the role as a catalog of logos and rewards people who can read a constraint and pick the engine, the store, and the reservation that match it.
There is no required prior certification. The recommended experience is the gap check. If projects, Identity and Access Management roles, Cloud Storage buckets, and gcloud are still a catalog rather than weekly work, sit Associate Cloud Engineer first. The data-engineer sitting assumes those controls and then asks which pipeline and which store to recommend. If the goal is the language of Google Cloud rather than the design of a data platform, Cloud Digital Leader is the closer match. Service definitions, shared responsibility, and product families belong there. They appear here as the vocabulary inside a pipeline, not as the whole job.
Skip it if the work you actually want is deploying someone else's design day to day without owning the transform, the warehouse model, or the slot reservation. That job is Associate Cloud Engineer. Skip it if the work you actually want is the compute shape, the network path, and a named case-study constraint across an entire solution. Those decisions are the Professional Cloud Architect sitting. Data products appear there as one recommendation among many. They are scored here as the entire job. Skip it if the stem you actually want is a custom training loop, a serving graph, and a model-operations program. Feature tables, BigQuery ML, embeddings, and retrieval-augmented generation appear on this outline as data preparation. They are not the Machine Learning Engineer sitting.
The current professional data-engineer credential Google publishes for this role is this one. Use this page for the Professional Data Engineer task list. Use the Google Associate Cloud Engineer exam guide when the stem is an operator control that happens to sit on a pipeline. Use the Google Cloud Digital Leader exam guide when the stem is still vocabulary. Use the Google Professional Cloud Architect exam guide when the stem is a solution recommendation that happens to include a data store.
Exam shape and how a pass is decided
The current exam guide publishes five sections and marks each weight as an approximation.
| Section | Weight | What gets tested |
|---|---|---|
| Designing data processing systems | about 22% | IAM and organization policies, encryption and keys, PII, data sovereignty, legal compliance, project and dataset and table architecture, multi-environment work, Dataform and Dataflow and Cloud Data Fusion, LLM query generation, pipeline monitoring, disaster recovery, ACID against availability, validation, portability and residency, migrations with BigQuery Data Transfer Service, Database Migration Service, Transfer Appliance, networking, and Datastream |
| Ingesting and processing the data | about 25% | Sources and sinks, transformation and orchestration logic, networking, encryption, cleansing, Dataflow, Apache Beam, Dataproc, Cloud Data Fusion, BigQuery, Pub/Sub, Apache Spark, the Hadoop ecosystem, Apache Kafka, batch against streaming including windowing and late data, processing logic, AI enrichment, Cloud Composer, Workflows, and CI/CD |
| Storing the data | about 20% | Access patterns, BigQuery, BigLake, AlloyDB, Bigtable, Spanner, Cloud SQL, Cloud Storage, Firestore, Memorystore, cost and performance, lifecycle, warehouse modeling and normalization, lake discovery and access and cost, Dataplex, Dataplex Catalog, and federated governance |
| Preparing and using data for analysis | about 15% | Visualization connectors, precalculated fields, BI Engine, materialized views, query troubleshooting, masking, IAM, Cloud Data Loss Prevention, BigQuery ML, embeddings and retrieval-augmented generation, sharing rules, published datasets and reports, and BigQuery sharing |
| Maintaining and automating data workloads | about 18% | Cost against business need, capacity for critical jobs, persistent against job-based Dataproc clusters, Composer DAGs, repeatable scheduling, BigQuery Editions and reservations, interactive against batch query jobs, Cloud Monitoring, Cloud Logging, the BigQuery admin panel, billing and quotas, multi-region or multi-zone jobs, corruption and missing data, and Cloud SQL or Redis failover |
Do the arithmetic on those approximations before building a plan. The five midpoints total 100%. Google still prints a tilde on every band, so no exact split exists to memorize. Section 2 sits alone at the top. Section 1 sits next. Section 3 sits in the middle. Section 5 sits next. Section 4 is the lightest band. Any study plan that gives five equal weeks overweights analysis and underweights ingest.
The logistics come off the certification page. The standard sitting is 2 hours. The format is 40 to 50 multiple choice and multiple select questions. Languages are English and Japanese. The registration fee is 200 USD, and tax where applicable. Delivery is online-proctored from a remote location or onsite-proctored at a testing center. Prerequisites are none. Validity for the professional certification is 2 years.
Case studies are not a scored share on this sitting. The certification page, the exam guide HTML, and the official exam guide PDF name no companies and no case-study percentage. Do not study a fifth company a dump site invented. Do not carry Professional Cloud Architect case-study habits onto this timer as if they were official here.
Scoring is the part most third-party pages get wrong. The certification page and the exam guide do not publish a numeric passing score, a scaled range, or a percent hedge. Google Cloud Certification Exam Policies and Exam Terms and Conditions say that if you pass an Exam, you will receive a digital certificate after Google has validated your score. That is the official pass language. This guide does not invent a 700 mark or a 70 percent story for a vendor that did not publish one.
The same terms page is the source for retakes and for how long the credential lasts. Associate and Professional exams allow a maximum of four attempts in a two year period. After a failed attempt, the wait is 14 days. After a second failed attempt, the wait is 60 days. After a third failed attempt, the wait is 365 days before a fourth attempt. Each attempt requires payment. A Professional Certification is valid for two years from the date of issue. Passing a Professional exam during renewal extends validity for two years from the date of passing.
Renewal has three official doors. The standard 2 hour exam. A 1 hour renewal exam of 20 questions, at 100 USD and tax where applicable. Designated courses or skill badges in Google Skills, which the certification page lists with a 1 year validity. Professional renewal eligibility on the terms page begins 60 days before expiration. After that window, the path back is the standard exam.
Two more details change how the sitting is taken. Google may update exam content at any time to reflect changes to Google Cloud technology. The certifications hub says exams are being updated for product updates announced at Google Cloud Next '26, including Gemini Enterprise Agent Platform and Google Cloud's data and analytics stack. The product names that matter on the timer are the names on the current exam guide, not the names on last month's console banner.
What the current outline changed
The exam guide does not publish a Microsoft-style change-log table. The official delta is the set of names the current pages print.
The certification page opens with a banner. The exam will soon be updated to reflect recent branding changes. The exam guide is the place to review the product names used on the exam. That sentence is the reason notes written only against today's console labels are already a risk. The current guide still says Dataproc, Cloud Composer, Cloud Data Loss Prevention, Dataplex, and Dataplex Catalog.
The certifications hub is more specific about the product wave. Exams are being updated to reflect product updates announced at Google Cloud Next '26, including Gemini Enterprise Agent Platform and Google Cloud's data and analytics stack. The current Professional Data Engineer guide already names Dataform, Datastream, BigLake, AlloyDB, BigQuery ML, embeddings and retrieval-augmented generation, BigQuery sharing, BigQuery Editions, and reservations. A study plan that treats generative AI as a side topic is studying last year's job.
Current product docs have already moved some of those names. The sitting has not finished that move.
| Name on the exam guide | Name on current first-party docs | What that means on the sitting |
|---|---|---|
| Dataproc | Managed Service for Apache Spark, the new name for Dataproc on Compute Engine and for Serverless for Apache Spark | A Spark or Hadoop cluster stem is still Dataproc on the outline |
| Cloud Composer | Managed Service for Apache Airflow. Cloud Composer is evolving into that name | An Airflow DAG stem is still Cloud Composer on the outline |
| Cloud Data Loss Prevention (Cloud DLP) | Sensitive Data Protection. Cloud DLP is now part of that product. The API is still the DLP API | A discover, classify, or de-identify stem is still Cloud DLP on the outline |
| Dataplex and Dataplex Catalog | Knowledge Catalog, the name as of 10 April 2026. The API, CLI, and IAM names are unchanged | A catalog, lineage, or federated-governance stem is still Dataplex on the outline |
| BigQuery sharing (Analytics Hub) | BigQuery sharing, formerly Analytics Hub | The exam guide already prints both names. A share-without-copy stem is BigQuery sharing |
Validity and renewal changed the calendar around the sitting even when the five section names stayed close. Professional Data Engineer lasts 2 years. A 1 hour renewal exam and a Google Skills course path sit beside the standard exam. Foundational and Associate credentials last 3 years on the same terms page. Mixing those clocks is how people schedule the wrong renewal door.
Three things follow for anyone holding older material.
The five section names are still the spine. Design, ingest, store, analyze, and maintain are still the scored map. Notes organized on those five headings are still structurally useful.
The product names inside those headings are in motion. Dataproc, Cloud Composer, Cloud DLP, and Dataplex are the names the exam guide still prints. A flashcard that only knows the new console label, and cannot map it back to the outline, is already off the branding banner.
This sitting has no official case-study PDFs. Time spent on invented company names is time taken from section 2.
How a processing engine gets picked
Section 2 is about 25%, the heaviest band, and the same discriminator shows up in design stems, migration stems, and operations stems as often as in pipeline ones. Section 1 then asks the candidate to secure and migrate the shape they just recommended.

The exam guide maps pipeline work to Dataflow, Apache Beam, Dataproc, Cloud Data Fusion, BigQuery, Pub/Sub, Apache Spark, the Hadoop ecosystem, and Apache Kafka. Read the control the stem is buying before reading the logo.
Pub/Sub is the ingest bus. What is Pub/Sub describes an asynchronous messaging service that decouples producers from processors, with latencies typically on the order of 100 milliseconds. Publishers send events. Subscribers process them. A common documented path is events into Pub/Sub, then Dataflow into BigQuery, Bigtable, or Cloud Storage. Pub/Sub is service-to-service. A stem about fan-out, decoupling, or a stream that several pipelines must read is Pub/Sub. A stem about transforming those events is not finished when the topic exists.
Dataflow is the unified batch and stream processor. Dataflow overview says the service reads from one or more sources, transforms the data, and writes to a destination. The programming model is the same for batch and stream. Default processing is exactly-once for every record. At-least-once mode is the documented option when duplicates are acceptable and the stem wants lower cost or lower latency. Google allocates worker virtual machines and deletes them when the job completes or is cancelled. Dataflow runs Apache Beam pipelines in Java, Python, or Go. A developer can write Beam code, deploy a template, or iterate in a notebook. A stem about windowing, late-arriving data, exactly-once streaming into BigQuery, or a portable Beam pipeline is Dataflow. Putting that work on a cluster the team will keep all month is the trap when the stem never asked for Spark.
Dataproc is the Spark and Hadoop cluster path. The current product docs call it Managed Service for Apache Spark and say that name replaces Dataproc on Compute Engine. The exam guide still says Dataproc. The clusters overview describes open-source batch, query, streaming, and machine-learning tools, clusters that start and shut down in 90 seconds or less on average, and built-in connections to BigQuery, Cloud Storage, Bigtable, Cloud Logging, and Cloud Monitoring. Hive, Pig, HBase, Flink, and Spark jobs still live here. Section 5.1 later scores persistent clusters against job-based clusters by name. A stem about an existing Spark job, a Hadoop ecosystem tool, or a cluster the team will size and then turn off is Dataproc. Rewriting that job in Beam because Dataflow is on this week's flashcards is the trap when the constraint is the existing Spark code.
Cloud Data Fusion is the visual integration path. Cloud Data Fusion overview describes a fully managed enterprise data-integration service powered by CDAP. Pipelines are a directed graph of source, transform, and sink nodes in Studio. Wrangler is the documented preparation surface. Execution often lands on a Dataproc provisioner. The visual design is still Data Fusion. A stem about a plugin, a Studio canvas, or a team that must not write Beam or Spark is Cloud Data Fusion. Naming the Dataproc cluster underneath as the whole answer is the trap when the stem asked for the integration UI.
Dataform is the SQL graph already inside BigQuery. Dataform overview describes ELT after raw data is already loaded. Analysts write SQLX, declare dependencies, run assertions, and compile Dataform core into Standard SQL. Compilation is hermetic. There is no internet during compile. A stem about tested, versioned SQL models in the warehouse is Dataform. A stem about streaming events, windowing, or late data is Dataflow.
Apache Kafka is on the outline as a service to identify. Managed Service for Apache Kafka is the current Google Cloud product for that engine. Pub/Sub is still the Google-native bus. A stem that already has Kafka partitions, offsets, or a Kafka migration is Kafka. A stem that wants a Google-native topic and a Dataflow subscriber is Pub/Sub.
| Engine | What the stem is buying | First Google Cloud product | What it does not do well |
|---|---|---|---|
| Ingest bus | Decouple producers from processors, fan-out, 100 millisecond class delivery | Pub/Sub | Transforming the payload |
| Unified batch and stream | Windowing, late data, exactly-once, Beam code or a template, workers that vanish | Dataflow | Keeping a Hadoop cluster API |
| Spark or Hadoop cluster | Existing Spark, Hive, HBase, or a cluster the team will size | Dataproc | A visual, no-code integration canvas |
| Visual integration | Studio graph, plugins, Wrangler, CDAP | Cloud Data Fusion | Replacing Beam when the team already owns the pipeline code |
| Warehouse SQL graph | ELT models, assertions, Git, already in BigQuery | Dataform | Streaming windowing |
| Kafka bus | Existing Kafka, offsets, partitions | Managed Service for Apache Kafka, when the stem names Kafka | A default Google-native topic |
Read the constraint the stem is buying. Events that must fan out before anyone transforms them are Pub/Sub. A portable Beam pipeline that must run batch and stream the same way is Dataflow. An existing Spark or Hadoop job is Dataproc. A visual canvas for a team that will not write code is Cloud Data Fusion. A tested SQL model already in BigQuery is Dataform. A Kafka estate that must stay Kafka is Kafka.
How a data store gets picked
Storage sits inside section 3 at about 20%, and again in the migration bullets of section 1 and the failover bullets of section 5. Most of those questions reduce to one skill. Name whether the stem is about a warehouse, a lake file, a wide-column operational store, a regional relational engine, a global relational ledger, a document, or a cache. Then name the service that implements that model.

Design an optimal storage strategy for your cloud workload still sorts Google Cloud storage into block, file, and object. The exam guide's own storage bullet is more specific for this sitting. It names BigQuery, BigLake, AlloyDB, Bigtable, Spanner, Cloud SQL, Cloud Storage, Firestore, and Memorystore. Read the access pattern before reading the brand.
| Model | What it is built for | First Google Cloud services to consider | What it does not do well |
|---|---|---|---|
| Warehouse | Interactive and batch SQL after the fact | BigQuery, with editions, reservations, BI Engine, and materialized views | The system of record for a checkout ledger |
| Lake, then query in place | Structured files that must stay in object storage | Cloud Storage with BigLake, Dataplex for discovery | Multi-row OLTP transactions |
| Wide-column operational | High throughput, single key, low latency, time series | Bigtable | Complex multi-row SQL as the system of record |
| Relational, regional | MySQL, PostgreSQL, or SQL Server with managed backups | Cloud SQL | Horizontal write scale across regions without a redesign |
| Relational, PostgreSQL HTAP | Analytical queries on live PostgreSQL transactions | AlloyDB | A cheap lift of SQL Server |
| Relational, global | Horizontal writes, strong consistency, multi-region | Spanner | A cheap lift of a single-zone SQL Server with no rewrite budget |
| Document | Flexible JSON, mobile clients | Firestore | Warehouse-scale analytics |
| Cache | Sub-millisecond keys, session state | Memorystore | A durable system of record |
| Object landing zone | Raw files, backups, media | Cloud Storage classes | Multi-row transactions |
BigQuery is the warehouse. Compute and storage are separate. Datasets, tables, partitions, and clusters are the documented shape. Load, export, query, and copy run as jobs. A stem about interactive SQL, a scheduled batch query, a BI dashboard, or a model trained with SQL is BigQuery. A stem about a checkout that must take writes in two regions with strong consistency is not.
BigLake is how BigQuery queries structured files that stay in an object store. Access delegation uses a service account to reach Cloud Storage, Amazon S3, or Azure Blob Storage. Users get table-level, row-level, and column-level security. Cloud Storage BigLake tables also support dynamic data masking. A stem about querying Parquet in a bucket without loading it, and still enforcing column security, is BigLake. Creating an external table with no delegation, then granting the user the bucket, is the trap when the stem asked for table-level control.
Cloud Storage is the object landing zone and the lake files themselves. Storage classes pick the cost curve. Standard has no minimum duration and no retrieval fee, and it is the class for data that is read often. Nearline has a 30 day minimum and a retrieval fee, and it is the class for data read about once a month. Coldline has a 90 day minimum and is the class for data read about once a quarter. Archive has a 365 day minimum, is the lowest-cost class, and still returns data in milliseconds. Autoclass is the switch when access frequency is unknown. Rapid storage is the zonal class for intensive reads and writes and only exists on Rapid Bucket. A stem that wants last year's logs to stay cheap but still open quickly is Archive, not a Bigtable table. A stem that says "cold" and "must wait hours" is not describing Cloud Storage Archive.
Bigtable is the wide-column operational store. It is a sparsely populated table that scales to billions of rows and thousands of columns, and to terabytes or petabytes. It is built for large amounts of single-keyed data with low latency and high read and write throughput. The HBase API is a documented client path. Time series, IoT, financial ticks, and marketing histories are the named shapes. Values are typically no larger than 10 MB. A stem about a row key, a column family, or a high-throughput time series is Bigtable. Putting that series in BigQuery because the analyst wants SQL later is the trap when the stem is still the operational write path.
Cloud SQL is the managed MySQL, PostgreSQL, or SQL Server service. Cloud SQL handles backups, high availability and failover, encryption, connectivity, storage, export and import, replication, maintenance, monitoring, and logging. Section 5.5 names Cloud SQL replication and failover again. A stem about lifting those engines with the least rewrite is Cloud SQL. A stem that names a SQL Server feature the team refuses to give up is still Cloud SQL, not Spanner.
AlloyDB is the managed PostgreSQL-compatible service built for demanding applications and for hybrid transactional and analytical processing. The product page says it runs complex analytical queries against live transactional data. AlloyDB AI puts vector search in the engine. A stem about PostgreSQL that must serve analytics on the live ledger, or a vector search next to that ledger, is AlloyDB. A stem about SQL Server is not.
Spanner is the globally distributed relational database with horizontal scale and strong consistency. A stem about a ledger that must take writes in more than one region without the team sharding by hand is Spanner. A stem about a regional reporting database that already runs on PostgreSQL and must move this quarter is Cloud SQL or AlloyDB, depending on the analytical load.
Firestore is the document model. Memorystore is the cache. Section 5.5 names Redis clusters for replication and failover. A stem about a mobile client that syncs documents is Firestore. A stem about a sub-millisecond session store is Memorystore. JSON in Cloud Storage does not make that bucket a document database.
Dataplex and Dataplex Catalog are the exam-guide names for the lake platform and the catalog. Current docs call the catalog Knowledge Catalog as of 10 April 2026. Section 3.3 scores lake discovery, access, and cost controls. Section 3.4 scores a data platform built with Dataplex, Dataplex Catalog, BigQuery, and Cloud Storage, and a federated governance model for distributed data. A stem about finding and governing files across buckets and warehouses is Dataplex. Querying those files is BigLake or BigQuery. Storing them is Cloud Storage.
Migrations have their own products in section 1.4. Datastream is the serverless change-data-capture path from operational databases into BigQuery, with Cloud Storage as a documented alternative sink. Sources include MySQL, Oracle, PostgreSQL, AlloyDB, SQL Server, MongoDB, and Spanner. BigQuery Data Transfer Service is the scheduled warehouse feed. Database Migration Service is the database cutover path. Transfer Appliance is the physical path when the dataset is too large for the network. A stem about ongoing CDC into the warehouse is Datastream. A stem about a nightly SaaS extract is BigQuery Data Transfer Service. A stem about moving MySQL to Cloud SQL is Database Migration Service. A stem about a truckload of files is Transfer Appliance.
How orchestration and capacity get picked
Section 5 is about 18%. Section 2.3 already named Cloud Composer and Workflows. Section 5.1 names persistent against job-based Dataproc clusters. Section 5.3 names BigQuery Editions, reservations, and interactive against batch query jobs. Those are one skill with three surfaces.

Cloud Composer is the Airflow path. The current product docs call it Managed Service for Apache Airflow and say Cloud Composer is evolving into that name. The exam guide still says Cloud Composer. The overview describes a fully managed orchestration service that creates, schedules, monitors, and manages pipelines across clouds and on-premises data centers. Workflows are Python DAGs. Environments are Airflow deployments on Google Kubernetes Engine. A Cloud Storage bucket holds the DAGs, logs, plugins, and data. Section 5.2 scores creating directed acyclic graphs for Cloud Composer by name. A stem about an Airflow DAG, a data job that must run after another data job, or an operator that launches Dataflow or Dataproc is Cloud Composer.
Workflows is the serverless service-orchestration path. A workflow is a YAML or JSON definition that executes Cloud Run, Cloud Run functions, Google Cloud services, and HTTP APIs in order. The service scales up as needed and incurs no charges while idle. A workflow can hold state, retry, poll, or wait for up to a year. A stem about a short HTTP sequence, a human callback, or an orchestration that must cost nothing when idle is Workflows. Putting that sequence on a Composer environment because the team already pays for Airflow is the trap when the stem never asked for a DAG.
Dataproc capacity is the cluster question hiding inside operations. A persistent cluster is the answer when the stem needs a warm Hadoop or Spark estate, a Hive metastore, or interactive notebooks on the cluster. A job-based cluster is the answer when the stem wants the cluster to exist for the job and then disappear. The clusters overview makes turning clusters off a first-class cost control. Section 5.1 writes that choice in the outline. A stem that keeps a cluster up all month for one nightly job is paying for idle nodes the exam is scoring.
BigQuery capacity is a different reservation. Understand BigQuery editions publishes three editions, and an on-demand model that can sit beside them on a per-project basis. Editions are a property of compute, not of storage. You can query a dataset as long as the edition supports the feature the query needs.
| Capacity choice | How it meters | What it includes on this outline | What it does not buy |
|---|---|---|---|
| On-demand | Bytes processed, up to 2,000 slots per project | Pay per query, BI Engine, most analysis features | Predictable slot baselines, capacity commitments |
| Standard edition | Slot-hours | Autoscaling reservations, QUERY and PIPELINE assignments | BI Engine, creating materialized views, continuous queries, Assured Workloads |
| Enterprise edition | Slot-hours, optional 1 year or 3 year commitments | BI Engine, idle capacity sharing, create materialized views, continuous queries, Omni | Managed disaster recovery, Assured Workloads |
| Highest edition | Slot-hours, optional commitments | Assured Workloads, managed disaster recovery, the Enterprise analysis set | A reason to pick it when the stem never named compliance or disaster recovery |
Introduction to workload management splits the same decision another way. On-demand charges for the data a query scans. Capacity-based billing allocates slots in reservations, with a 50 slot minimum per reservation, and can autoscale. Interactive query jobs are the path when a person is waiting on the result. Batch query jobs are the path when the work can wait for idle slots. Section 5.3 writes that pair on the outline. A stem about a dashboard that must return while someone watches is interactive, and often wants BI Engine memory on top of slots. A stem about a 2 a.m. transform that can wait is batch.
BI Engine is an in-memory acceleration layer that caches the parts of columns and partitions a dashboard actually reads. Reservations are in GiB of memory and are managed separately from slot reservations. Standard edition does not include BI Engine. Enterprise, the highest edition, and on-demand do. Preferred tables are how a team pins acceleration to the dashboard tables. A stem about a slow Looker or Tableau dashboard on a hot subset is BI Engine. Buying more slots and ignoring memory is the trap when the outline already named BI Engine.
Materialized views precompute the join or aggregation the dashboard keeps repeating. Standard edition can query a materialized view that already exists. Creating and refreshing one is an Enterprise feature on the editions table. A stem about precalculated fields in section 4.1 often lands here or on a scheduled Dataform model.
Sharing is its own product. BigQuery sharing, formerly Analytics Hub, publishes listings in a data exchange so subscribers get a linked dataset without a copy of the data. Publishers can restrict egress. Subscribers query in place. A stem about sharing a curated dataset with another organization, without replicating it, is BigQuery sharing. Exporting a CSV and mailing it is the trap.
BigQuery ML trains and runs models with GoogleSQL inside the warehouse. Models live in datasets. SQL practitioners do not have to move the table to a notebook to get a forecast, a cluster, or a contribution analysis. The same surface reaches Gemini Enterprise Agent Platform models and Cloud AI APIs. Section 4.2 also names embeddings and retrieval-augmented generation for unstructured data. A stem about preparing features, training in SQL, or building embeddings for retrieval is this sitting. A stem about a custom training loop on GPUs is the Machine Learning Engineer outline.
Sensitive Data Protection is the current name for the product that still appears on the exam guide as Cloud Data Loss Prevention. The API is still the DLP API. The job is to discover, classify, and de-identify sensitive data. A stem about masking a column before a dashboard, or inspecting a bucket for personal data, is Cloud DLP on the outline.
How to study the current blueprint
Order the weeks by the weight bands, not by the order the products appear in a catalog.
Week 1. Ingesting and processing, the band near 25%. Publish one event stream to a Pub/Sub topic and write it to BigQuery two ways. First, run a Dataflow template from Pub/Sub to BigQuery and confirm the job's workers disappear when you stop it. Second, leave the same topic in place and write down why a second subscriber can read the same events without a second ingest. Submit one Apache Beam pipeline that windows late data, then turn on at-least-once mode and write down which guarantee you gave up. Create a Dataproc cluster, submit a Spark job, delete the cluster, then create a job-based cluster that exists only for that job. Open Cloud Data Fusion Studio long enough to build a three-node pipeline and write down that the Dataproc provisioner underneath is not the product the stem named. Load a raw table into BigQuery and turn it into a Dataform model with one assertion. Finish the week by naming Kafka only when the stem already has Kafka.
Week 2. Designing data processing systems, the band near 22%. Build the governance story the outline actually prints. Create two projects under one folder, one for development and one for production, and assign an IAM role and an organization policy at the folder so both inherit. Create a dataset and a table, then write down who owns the project, who owns the dataset, and who owns the table. Store a customer-managed key and encrypt one BigQuery table or one Cloud Storage bucket with it. Turn on a regional constraint and write down why a multi-region dataset would fight data sovereignty. Run Sensitive Data Protection, the product the exam guide still calls Cloud DLP, against a sample table and mask one column. Then run the four migration products far enough to tell them apart. Create a Datastream stream from a supported operational database toward BigQuery. Open BigQuery Data Transfer Service and name one scheduled source. Open Database Migration Service and name a Cloud SQL or AlloyDB target. Open Transfer Appliance long enough to say it is the physical path. Finish the week on Dataform, Dataflow, and Cloud Data Fusion as cleaning tools, the way section 1.2 lists them, and on one LLM-generated query that you still validate.
Week 3. Storing the data, the band near 20%. Create one object, one warehouse table, and one operational key. Put the same CSV in Cloud Storage as Standard, Nearline, Coldline, and Archive, and write down the minimum duration and the retrieval fee for each class. Create a BigLake table over the Standard object and grant a user the table without granting the bucket. Create a BigQuery table that a dashboard will scan, partition it by date, and cluster it by the filter the dashboard actually uses. Create Cloud SQL for one engine the outline still names, then open the AlloyDB overview and write down the constraint that would have forced PostgreSQL HTAP instead. Open the Spanner product page and write down the constraint that would have forced global writes. Create a Bigtable table with one row key and one column family, and write a time-series row. Create a Firestore document and a Memorystore instance and write down why neither is the warehouse. Open Dataplex, the catalog the docs now call Knowledge Catalog, and register the bucket and the dataset so discovery is not a spreadsheet.
Week 4. Operations near 18%, analysis near 15%, and a full review. Write a Cloud Composer DAG that launches the Dataflow job from week 1 and then the Dataform run from week 1. Write a Workflows definition that calls one HTTP endpoint and one BigQuery job, and write down why that definition incurs no charge while idle. Create a BigQuery reservation in Standard edition and one in Enterprise edition, assign a project to each, and run the same dashboard query both ways so BI Engine and materialized-view creation are not theoretical. Submit one interactive query and one batch query and watch where they queue. Create a Cloud Monitoring dashboard and a log-based alert on a failed Dataflow job. Fail over a Cloud SQL high-availability instance in a test project, or walk the documented failover path if you will not break the instance. Train one BigQuery ML model with SQL. Generate embeddings for a small unstructured table and write down the retrieval step that section 4.2 is scoring. Publish one dataset through BigQuery sharing and subscribe from a second project so the linked dataset is a pointer, not a copy. Then sit with the exam guide and, for each bullet, name the product, the model, and the constraint that would have forced that product.
Take any official sample questions Google still links from the certification page and walk the exam tutorial at least once before using this outline against a clock. The sample set exists so the question format costs nothing on the timer.
Traps that look like easy elimination
Data-engineer items rarely give one plausible answer and three absurd ones. They give two controls that both sound correct and one detail that picks between them.
- There is no prerequisite certification. Recommended experience is still 3 or more years, including 1 or more years on Google Cloud. Showing up with only vocabulary from Cloud Digital Leader is how section 2 is lost.
- The standard sitting is 40 to 50 questions in 2 hours. Notes that quote a 50 to 60 question professional sitting are quoting a different exam.
- Google does not publish a numeric passing score on the certification page or the exam guide. A third-party 700 or 70 percent figure is not official language for this sitting.
- This sitting has no official case-study PDFs. A company name that does not appear on the exam guide is not part of the outline.
- Pub/Sub moves messages. Dataflow transforms them. A topic is not a pipeline.
- Dataflow is unified batch and stream and defaults to exactly-once. At-least-once is a cost and latency trade, not a silent default.
- Dataproc is the exam-guide name for the Spark and Hadoop cluster. Managed Service for Apache Spark is the current docs name. The outline still says Dataproc.
- A Dataproc provisioner under Cloud Data Fusion does not turn a Studio pipeline into a Dataproc answer.
- Dataform is ELT inside BigQuery. It is not a streaming engine.
- Kafka is the answer when the stem already has Kafka. Pub/Sub is the Google-native bus.
- Cloud Composer is the exam-guide name for Airflow DAGs. Managed Service for Apache Airflow is the current docs name. The outline still says Cloud Composer.
- Workflows is serverless service orchestration with no idle charge. An Airflow environment is not that product.
- A persistent Dataproc cluster is for a warm estate. A job-based cluster is for a job that should disappear. Keeping the cluster up for one nightly job is the cost trap section 5.1 scores.
- BigQuery is the warehouse. Bigtable is the operational wide-column store. Cloud SQL is the managed engine lift. AlloyDB is PostgreSQL HTAP. Spanner is global relational. Firestore is documents. Memorystore is cache. Cloud Storage is objects.
- BigLake queries files in place with access delegation. An external table that also requires bucket access is a weaker control when the stem asked for table-level security.
- Standard, Nearline, Coldline, and Archive are duration and access trades, not durability trades. Archive still opens in milliseconds.
- Datastream is CDC. BigQuery Data Transfer Service is a scheduled warehouse feed. Database Migration Service is a database cutover. Transfer Appliance is a physical ship. Those are four answers.
- On-demand BigQuery charges for bytes scanned. Capacity charges for slots. Interactive jobs wait with the person. Batch jobs wait for idle slots.
- BI Engine is memory, in GiB, separate from slots. Standard edition does not include it.
- Materialized views that already exist can be queried on Standard. Creating them is an Enterprise feature on the editions table.
- BigQuery sharing publishes a listing. The subscriber gets a linked dataset, not a copied table.
- Cloud DLP is the exam-guide name. Sensitive Data Protection is the current docs name. The API is still the DLP API.
- Dataplex and Dataplex Catalog are the exam-guide names. Knowledge Catalog is the current docs name as of 10 April 2026.
- BigQuery ML trains in SQL inside the warehouse. A custom GPU training loop is a different sitting.
- Embeddings and retrieval-augmented generation are data-preparation bullets in section 4.2. They are not a reason to ignore the pipeline sections.
If deleting the scenario still lets you pick the answer from the service name, the question is easier than the live exam.
How this maps to CloudFluently
Start with the official material. The Professional Data Engineer certification page carries the audience profile, the 2 hour timer, the 40 to 50 question format, the 200 USD fee, the 2 year validity, and the three renewal doors. The Professional Data Engineer exam guide carries the five sections and the product names. The official renewal outline is the Professional Data Engineer renewal exam guide PDF. Exam Terms and Conditions explain what it means to pass an Exam, how long a Professional Certification lasts, and the 14 day, 60 day, and 365 day retake waits.
Professional Data Engineer sits on top of Google Cloud vocabulary and Google Cloud operations, and that ground is already live here. Shared responsibility, product families, and the idea of a cloud bill are the Cloud Digital Leader study notes and the Cloud Digital Leader practice exam set. The GCP Digital Leader roadmap sequences that ground, and the Cloud Digital Leader exam guide is the first companion page. Projects, IAM, Compute Engine, Cloud Storage, VPC networks, and the operator form of BigQuery and Pub/Sub are the Associate Cloud Engineer study notes and the Associate Cloud Engineer practice exam sets. The GCP Associate Cloud Engineer roadmap sequences that work, and the Associate Cloud Engineer exam guide is the operator companion.
The architect reading of a data store, when the stem is still a solution recommendation rather than a pipeline, is the Professional Cloud Architect exam guide.
Work those until the vocabulary, the operator controls, and the architect habit are automatic, then use this page and the official exam guide for the data-engineer-only skills: the processing engine, the store model, the scheduler, and the reservation that holds them.
Frequently Asked Questions
What is the passing score for Professional Data Engineer? Google does not publish a numeric passing score on the certification page or the exam guide. Exam Terms and Conditions say that if you pass an Exam, you receive a digital certificate after Google has validated your score.
How long is the exam and how many questions are there? The standard sitting is 2 hours with 40 to 50 multiple choice and multiple select questions. The renewal exam is 1 hour with 20 questions.
Does this exam use case studies? The certification page and the official exam guide do not publish case studies for this sitting. Study the five sections and the named products.
Is there a prerequisite? No. Google recommends 3 or more years of industry experience, including 1 or more years designing and managing solutions using Google Cloud.
How long does the certification last? A Professional Certification is valid for two years from the date of issue. Renewal is the standard exam, the 1 hour renewal exam, or designated Google Skills courses or skill badges. The Skills path is listed with a 1 year validity on the certification page.
Can I retake it if I fail? Associate and Professional exams allow four attempts in two years. The waits are 14 days after the first fail, 60 days after the second, and 365 days after the third. Each attempt is paid.
What changed on the current outline? The certification page says the exam will soon be updated for branding. The current guide still names Dataproc, Cloud Composer, Cloud DLP, Dataplex, and Dataplex Catalog. Current docs already print Managed Service for Apache Spark, Managed Service for Apache Airflow, Sensitive Data Protection, and Knowledge Catalog. The certifications hub points at Google Cloud Next '26 product updates, including the data and analytics stack.
Which section is heaviest? Ingesting and processing the data is about 25%. Designing data processing systems is about 22%. Storing the data is about 20%. Maintaining and automating data workloads is about 18%. Preparing and using data for analysis is about 15%.
Does Google publish a 700 or 70 percent pass mark? No. That figure is not on the certification page, the exam guide, or the terms page opened for this guide.
How does this exam relate to Associate Cloud Engineer, Cloud Digital Leader, and Professional Cloud Architect? Cloud Digital Leader scores vocabulary. Associate Cloud Engineer scores the operator controls. Professional Cloud Architect scores the solution recommendation. Professional Data Engineer scores the pipeline, the store, the scheduler, and the reservation. The official outline is the Professional Data Engineer exam guide. Use Google for the task statements. Use this page for how those statements get picked, what the current names are, and how to sequence the blueprint.
