Google Professional Data Engineer 1-10

表示モード
画像位置
文字位置
理解度の自動記録
STATUS FILTER

Choose confidence levels to display

Loading...
Q1Google Professional Data Engineer
Show answer
Correct answer: C. Dropout Methods
Fitting the training data well while performing poorly on new data is a classic sign of overfitting.
Dropout is a regularization technique that disables some neurons during training to prevent excessive reliance on specific features, improving generalization.
The standard technique to suppress overfitting and improve generalization is Dropout.
Threading and Serialization are processing and storage mechanisms and provide no remedy.
Dimensionality Reduction can reduce the number of features, but the standard technique that directly suppresses overfitting in a neural network is Dropout.
Machine Learning Crash Course: Overfitting (Google for Developers)
Q2Google Professional Data Engineer
Show answer
Correct answer: B. Continuously retrain the model by combining the existing data with the new data.
To keep the model accurate, you need to incorporate new trends while also preserving past patterns.
Retraining on only the new data tends to lose useful knowledge learned in the past, while using only the existing data cannot keep up with change.
To follow shifting preferences, retrain by combining the existing data with the new data.
This reflects the latest trends while maintaining stable predictions.
The point of this question is the composition of the training data, not the handling of a test set.
Best practices for ML development (Google Cloud)
Q3Google Professional Data Engineer
Show answer
Correct answer: C. Normalize the master patient records table into a patients table and a visits table, and create the other tables needed to avoid self-joins.
A single-table design that heavily uses self-joins sees join cost surge as the data grows, degrading performance.
Normalizing to avoid self-joins is the most reasonable design change to achieve both performance and scalability.
Normalizing the master table into a patients table and a visits table and arranging related tables eliminates self-joins and reduces query load.
Beefing up the server (A) is a stopgap and does not address the root cause.
Splitting by date or clinic (B and D) creates usage constraints and extra effort.
Best practices for database schema design (Google Cloud)
Q4Google Professional Data Engineer
Show answer
Correct answer: A. Edit the report settings to disable caching.
Data Studio (now Looker Studio) has a caching mechanism to speed up display, and this can cause the latest data not to be reflected.
To display the latest data, adjust the cache (data freshness) in the report settings.
The BigQuery-side cache (B) and browser operations (C and D) do not remedy this problem.
Note that the product was renamed from Data Studio to Looker Studio, inheriting its functionality, and the configuration concept is the same.
Manage data freshness (Looker Studio Help)
Q5Google Professional Data Engineer
Show answer
Correct answer: D. Run a Google Cloud Dataflow batch pipeline to import the data into BigQuery, and send errors to a separate dead-letter table for analysis.
For data that may contain malformed or corrupted rows, a design that validates at ingestion and isolates problem rows without halting processing is important.
To isolate bad rows without stopping, use Dataflow to route them to a dead-letter table.
Valid rows are loaded into BigQuery, and error rows can be investigated later.
Setting max_bad_records to 0 (C) fails the entire load if even one row is bad.
Checking with SQL (A) and monitoring alerts (B) are insufficient for isolating corrupted rows and continuing processing.
Dataflow pipeline best practices (Google Cloud)
Q6Google Professional Data Engineer
Show answer
Correct answer: B. Retry the query with exponential backoff, up to a maximum of 15 minutes.
In a large-scale system, concentrating short-interval retries during a database failure floods the recovering server with load and worsens the outage.
Retries during a failure should widen intervals with exponential backoff to avoid concentrated load.
Setting an upper limit lets you wait for automatic recovery while keeping server load down.
A restart command (A) is not the front end’s responsibility, retrying every second (C) causes overload, and fixing a lower frequency (D) lacks flexibility.
Retry strategy (Cloud Storage / Google Cloud)
Q7Google Professional Data Engineer
Show answer
Correct answer: A. Linear regression
Linear regression is well suited to a regression problem that predicts continuous values such as housing prices.
For predicting continuous values in a low-resource environment, lightweight linear regression is optimal.
Linear regression is computationally inexpensive and can train and infer efficiently even on a single resource-limited VM.
Logistic classification (B) is for classification and is unsuitable for continuous-value prediction.
An RNN (C) and a feedforward neural network (D) have high expressive power but require substantial compute resources and are not suitable for a constrained environment.
The CREATE MODEL statement (linear regression / BigQuery ML)
Q8Google Professional Data Engineer
Show answer
Correct answer: D. Use a ROW_NUMBER window function partitioned by the unique ID, together with a condition that the row number is 1 (WHERE row = 1).
With streaming inserts, the same row may be inserted more than once.
For deduplication, use ROW_NUMBER partitioned by the unique ID and keep only row number 1.
By partitioning by the unique ID with PARTITION BY and numbering each ID’s rows with ROW_NUMBER(), keeping only the rows where the row number is 1 yields one row per ID.
LIMIT 1 (A) returns only a single row overall, and GROUP BY + SUM (B) inappropriately aggregates the values.
LAG (C) is for comparing adjacent rows, and the standard technique for deduplication is ROW_NUMBER.
Analytic (window) function concepts (BigQuery / Google Cloud)
Q9Google Professional Data Engineer
Show answer
Correct answer: D. `bigquery-public-data.noaa_gsod.gsod*`
For wildcard tables, you need to append an asterisk (*) to the end of the table name to reference multiple tables together, and enclose the entire project and table name—which contains hyphens—in backticks (`).
A wildcard table appends * at the end and encloses the entire identifier in backticks.
D satisfies this syntax and can narrow down the target tables with _TABLE_SUFFIX.
Options without backticks (B) or with incorrect enclosure or quotation marks (A and C) cannot correctly interpret an identifier containing hyphens or the wildcard, and result in a syntax error.
Query multiple tables using a wildcard table (BigQuery / Google Cloud)
Q10Google Professional Data Engineer
Show answer
Correct answer: B, D, E
To achieve the principle of least privilege, measures that narrow the accessible scope itself are effective.
Achieving least privilege centers on three points: restricting by role, limiting API access, and separating data.
The applicable ones are restricting table access by role (B), limiting BigQuery API access to authorized users (D), and separating the data into multiple tables or databases (E).
Encryption (C) is effective for data protection but is a different concern from least privilege, disabling writes (A) does not narrow the read scope, and audit logging (F) is after-the-fact violation detection—none of these directly minimize access itself.
Introduction to access control (BigQuery / Google Cloud)