表示モード
画像位置
文字位置
理解度の自動記録
Q1Google Professional Data Engineer
Q1. Your company has built a neural network model with many neurons and layers using TensorFlow.
The model fits the training data well.
However, when tested against new data, its performance degrades.
Which technique can you adopt to address this?
Show answer
Correct answer: C. Dropout Methods
Fitting the training data well while performing poorly on new data is a classic sign of overfitting.Dropout is a regularization technique that disables some neurons during training to prevent excessive reliance on specific features, improving generalization.
The standard technique to suppress overfitting and improve generalization is Dropout.
Threading and Serialization are processing and storage mechanisms and provide no remedy.
Dimensionality Reduction can reduce the number of features, but the standard technique that directly suppresses overfitting in a neural network is Dropout.
Machine Learning Crash Course: Overfitting (Google for Developers)
Q2Google Professional Data Engineer
Q2. You are building a model that makes clothing recommendations.
Because you know that users’ fashion preferences are likely to change over time, you build a data pipeline that streams new data back to the model as soon as it becomes available.
How should you use this data to train the model?
Show answer
Correct answer: B. Continuously retrain the model by combining the existing data with the new data.
To keep the model accurate, you need to incorporate new trends while also preserving past patterns.Retraining on only the new data tends to lose useful knowledge learned in the past, while using only the existing data cannot keep up with change.
To follow shifting preferences, retrain by combining the existing data with the new data.
This reflects the latest trends while maintaining stable predictions.
The point of this question is the composition of the training data, not the handling of a test set.
Best practices for ML development (Google Cloud)
Q3Google Professional Data Engineer
Q3. As a pilot project, you designed a database of patient records for a few hundred patients across three clinics.
The design represented all patients and their visit history in a single table, and used a self-join to generate reports.
The server’s resource utilization was 50%.
The scope of the project then expanded, and the database now needs to store 100 times as many patient records.
Reports can no longer be run because they take too long or fail due to insufficient compute resources.
How should you adjust the database design?
Show answer
Correct answer: C. Normalize the master patient records table into a patients table and a visits table, and create the other tables needed to avoid self-joins.
A single-table design that heavily uses self-joins sees join cost surge as the data grows, degrading performance.Normalizing to avoid self-joins is the most reasonable design change to achieve both performance and scalability.
Normalizing the master table into a patients table and a visits table and arranging related tables eliminates self-joins and reduces query load.
Beefing up the server (A) is a stopgap and does not address the root cause.
Splitting by date or clinic (B and D) creates usage constraints and extra effort.
Best practices for database schema design (Google Cloud)
Q4Google Professional Data Engineer
Q4. You are building an important report in Google Data Studio 360 (now Looker Studio) for a large team.
The report uses Google BigQuery as its data source.
You notice that data less than one hour old does not appear in the visualizations.
What should you do?
Show answer
Correct answer: A. Edit the report settings to disable caching.
Data Studio (now Looker Studio) has a caching mechanism to speed up display, and this can cause the latest data not to be reflected.To display the latest data, adjust the cache (data freshness) in the report settings.
The BigQuery-side cache (B) and browser operations (C and D) do not remedy this problem.
Note that the product was renamed from Data Studio to Looker Studio, inheriting its functionality, and the configuration concept is the same.
Manage data freshness (Looker Studio Help)
Q5Google Professional Data Engineer
Q5. An external customer provides you with a daily dump of data from their database.
The data flows into Google Cloud Storage (GCS) as comma-separated (CSV) files.
You want to analyze this data in Google BigQuery, but it may contain malformed or corrupted rows.
How should you build this pipeline?
Show answer
Correct answer: D. Run a Google Cloud Dataflow batch pipeline to import the data into BigQuery, and send errors to a separate dead-letter table for analysis.
For data that may contain malformed or corrupted rows, a design that validates at ingestion and isolates problem rows without halting processing is important.To isolate bad rows without stopping, use Dataflow to route them to a dead-letter table.
Valid rows are loaded into BigQuery, and error rows can be investigated later.
Setting max_bad_records to 0 (C) fails the entire load if even one row is bad.
Checking with SQL (A) and monitoring alerts (B) are insufficient for isolating corrupted rows and continuing processing.
Dataflow pipeline best practices (Google Cloud)
Q6Google Professional Data Engineer
Q6. Your weather app queries a database every 15 minutes to obtain the current temperature.
The front end runs on Google App Engine and serves millions of users.
How should you design the front end to handle database failures?
Show answer
Correct answer: B. Retry the query with exponential backoff, up to a maximum of 15 minutes.
In a large-scale system, concentrating short-interval retries during a database failure floods the recovering server with load and worsens the outage.Retries during a failure should widen intervals with exponential backoff to avoid concentrated load.
Setting an upper limit lets you wait for automatic recovery while keeping server load down.
A restart command (A) is not the front end’s responsibility, retrying every second (C) causes overload, and fixing a lower frequency (D) lacks flexibility.
Retry strategy (Cloud Storage / Google Cloud)
Q7Google Professional Data Engineer
Q7. You are building a model to predict housing prices.
Due to budget constraints, it must run on a single, resource-limited virtual machine (VM).
Which learning algorithm should you use?
Show answer
Correct answer: A. Linear regression
Linear regression is well suited to a regression problem that predicts continuous values such as housing prices.For predicting continuous values in a low-resource environment, lightweight linear regression is optimal.
Linear regression is computationally inexpensive and can train and infer efficiently even on a single resource-limited VM.
Logistic classification (B) is for classification and is unsuitable for continuous-value prediction.
An RNN (C) and a feedforward neural network (D) have high expressive power but require substantial compute resources and are not suitable for a constrained environment.
The CREATE MODEL statement (linear regression / BigQuery ML)
Q8Google Professional Data Engineer
Q8. You are building a new real-time data warehouse for your company and using Google BigQuery streaming inserts.
There is no guarantee that data is sent exactly once, but each row has a unique ID and an event timestamp.
When querying the data interactively, you want to exclude duplicates.
Which query type should you use?
Show answer
Correct answer: D. Use a ROW_NUMBER window function partitioned by the unique ID, together with a condition that the row number is 1 (WHERE row = 1).
With streaming inserts, the same row may be inserted more than once.For deduplication, use ROW_NUMBER partitioned by the unique ID and keep only row number 1.
By partitioning by the unique ID with PARTITION BY and numbering each ID’s rows with ROW_NUMBER(), keeping only the rows where the row number is 1 yields one row per ID.
LIMIT 1 (A) returns only a single row overall, and GROUP BY + SUM (B) inappropriately aggregates the values.
LAG (C) is for comparing adjacent rows, and the standard technique for deduplication is ROW_NUMBER.
Analytic (window) function concepts (BigQuery / Google Cloud)
Q9Google Professional Data Engineer
Q9. Your company uses WILDCARD tables to query data across similarly named tables.
Currently, the SQL statement fails with the following error.
# Syntax error : Expected end of statement but got “-” at [4:11]
SELECT age FROM bigquery-public-data.noaa_gsod.gsod WHERE age != 99 AND _TABLE_SUFFIX = ‘1929’ ORDER BY age DESC
Which table name makes this SQL statement work correctly?
Show answer
Correct answer: D. `bigquery-public-data.noaa_gsod.gsod*`
For wildcard tables, you need to append an asterisk (*) to the end of the table name to reference multiple tables together, and enclose the entire project and table name—which contains hyphens—in backticks (`).A wildcard table appends * at the end and encloses the entire identifier in backticks.
D satisfies this syntax and can narrow down the target tables with _TABLE_SUFFIX.
Options without backticks (B) or with incorrect enclosure or quotation marks (A and C) cannot correctly interpret an identifier containing hyphens or the wildcard, and result in a syntax error.
Query multiple tables using a wildcard table (BigQuery / Google Cloud)
Q10Google Professional Data Engineer
Q10. Your company operates in a heavily regulated industry.
One requirement is that each user can access only the minimum information needed to do their job.
You want to enforce this requirement in Google BigQuery.
Choose three approaches you could take.
(Choose three)
Show answer
Correct answer: B, D, E
To achieve the principle of least privilege, measures that narrow the accessible scope itself are effective.Achieving least privilege centers on three points: restricting by role, limiting API access, and separating data.
The applicable ones are restricting table access by role (B), limiting BigQuery API access to authorized users (D), and separating the data into multiple tables or databases (E).
Encryption (C) is effective for data protection but is a different concern from least privilege, disabling writes (A) does not narrow the read scope, and audit logging (F) is after-the-fact violation detection—none of these directly minimize access itself.
Introduction to access control (BigQuery / Google Cloud)
