表示モード
画像位置
文字位置
理解度の自動記録
Q1Google Professional Machine Learning Engineer
Q1. You work as an ML researcher at a large insurance company and are experimenting with a large language model (LLM) called Gemini.
You plan to deploy this model for an internal use case, and you need to minimize inference time while retaining full control over the model’s underlying infrastructure.
Which serving configuration should you use for this task?
Show answer
Correct answer: B. Create your own custom YAML manifest and manually deploy the model to a Google Kubernetes Engine (GKE) cluster.
For a requirement that demands full control over the underlying infrastructure while minimizing inference latency, manually deploying to a GKE cluster with a custom YAML manifest is the best approach.GKE lets you fine-tune node types, accelerator placement, and scheduling in detail, so you can keep inference latency to a minimum.
Vertex AI’s one-click deploy (Option A) and deployment via Model Garden (Option D) are convenient but offer far less freedom in controlling the infrastructure.
Vertex AI’s custom inference container (Option C) allows some customization, but you cannot command the entire underlying platform as fully as with GKE.
Option B is correct because it achieves both maximum control and low latency.
Google Kubernetes Engine (GKE) Documentation
Q2Google Professional Machine Learning Engineer
Q2. You work for an organization that runs a streaming music service.
You operate a production custom model that recommends the “next song” based on a user’s recent listening history.
This model is already deployed to a Vertex AI endpoint.
Recently you retrained the same model on new data, and the offline test results were good.
You want to validate the new model in production while keeping complexity to a minimum.
What should you do?
Show answer
Correct answer: C. Deploy the new model to the existing Vertex AI endpoint. Use traffic splitting to send 5% of production traffic to the new model. Monitor end-user metrics such as listening time, and if the metrics improve compared with the other model, gradually increase the share of traffic to the new model.
Traffic splitting is a feature that lets you distribute prediction requests across multiple models or versions within the same Vertex AI endpoint.Because you can deploy the new model to the existing endpoint and route only 5% of traffic to it, you can perform a production canary test without creating a new endpoint or a separate service.
This minimizes complexity while allowing a gradual rollout as you watch metrics such as listening time.
Option A requires building a new endpoint and a routing service, which is complex.
Option B stays at offline evaluation and does not amount to production validation.
Option D is a wholesale switch that is high-risk and is not a gradual test.
Vertex AI: Model Deployment and Traffic Splitting
Q3Google Professional Machine Learning Engineer
Q3. Your organization has asked you to compare a variety of widely available ML models for a generative AI use case.
The models under comparison are also available on Google Cloud.
Each team has provided you with a curated internal benchmark dataset tailored to its specific use case or task.
You need to submit a comprehensive report summarizing your recommendations.
You want to evaluate the models in the most efficient way possible.
What should you do?
Show answer
Correct answer: A. Use Model Garden to deploy the candidate models to Vertex AI endpoints. Use the Gen AI Evaluation Service API to evaluate each model’s performance on the internal benchmark datasets. Report the best model based on the experiment results.
Vertex AI’s Gen AI Evaluation Service is a feature designed to automatically evaluate and compare multiple foundation models or LLMs in parallel using your own datasets.By using the evaluation service API integrated with Model Garden, you eliminate the need to write your own evaluation scripts or manage complex experiments, allowing you to compare the models most efficiently.
Because the question emphasizes using the “curated internal benchmark datasets,” Options B and C, which substitute open-source data, do not meet the requirement.
Option D requires downloading weights and writing scripts, which is labor-intensive and runs counter to the efficiency requirement.
Vertex AI: Gen AI Evaluation Service API
Q4Google Professional Machine Learning Engineer
Q4. As your company’s lead ML engineer, you are responsible for building an ML model that digitizes scanned customer forms.
You have developed a TensorFlow model that converts scanned images into text and stores it in Cloud Storage.
You need to apply this ML model to the data aggregated at the end of each business day, with as little manual intervention as possible.
What should you do?
Show answer
Correct answer: A. Use the batch prediction feature of AI Platform (now Vertex AI).
Batch prediction is a process that runs predictions on a large volume of data all at once, and it is well suited to cases that do not require real-time responsiveness and can be processed on a schedule.For a use case that digitizes a batch of forms at the end of each business day, Option A—submitting a batch prediction job against the input data in Cloud Storage—is optimal.
The results are also written to Cloud Storage, minimizing manual work and scaling automatically.
Options B, C, and D are geared toward real-time/online use and are either overkill or unsuitable for daily batch processing.
Note that AI Platform is now integrated into Vertex AI, but the concept of batch prediction and the correct answer remain valid today.
Vertex AI: Get Batch Predictions
Q5Google Professional Machine Learning Engineer
Q5. One of your models is trained on data supplied by a third-party data broker.
This data broker does not reliably notify you of changes to the data format.
You want to make your model’s training pipeline more robust against such problems.
What should you do?
Show answer
Correct answer: A. Use TensorFlow Data Validation (TFDV) to detect and flag schema anomalies.
TensorFlow Data Validation (TFDV) is a library for understanding, validating, and monitoring machine learning data.Because it can automatically detect and report schema anomalies such as missing features, new features, and data-type mismatches, it is ideal for hardening a pipeline against format changes.
It can also generate descriptive statistics and visualizations, and embedding it in the training pipeline lets you continuously ensure data quality.
Option B forcibly coerces values and thereby hides the problem.
Options C and D require custom implementations, lack comprehensiveness, and are inferior to TFDV for schema validation.
TensorFlow Data Validation (TFDV) Guide
Q6Google Professional Machine Learning Engineer
Q6. You are developing an ML model for CT-scan image segmentation using AI Platform (now Vertex AI).
You need to frequently update the model architecture based on the latest research papers and retrain it on the same dataset to benchmark its performance.
You want to keep code version control while minimizing compute cost and manual work.
What should you do?
Show answer
Correct answer: C. Integrate Cloud Build with Cloud Source Repositories so that a push of new code to the repository triggers retraining.
To keep code version control while minimizing compute cost and manual work, a configuration that integrates a source repository with CI and automatically launches retraining on a code push is appropriate.Option C—integrating Cloud Build with the repository and running retraining triggered by a push—is the optimal solution that satisfies both version control and automation.
Option A relies on coarse Storage monitoring for change detection, and Option B requires a manual run every time.
Option D’s daily schedule runs regardless of code changes, creating waste.
Note that Cloud Source Repositories stopped accepting new customers as of June 2024 (existing users can continue to use it), but the concept of triggering Cloud Build on a code push remains valid today.
Cloud Build: Creating and Managing Build Triggers
Q7Google Professional Machine Learning Engineer
Q7. You are building and operating a production system responsible for sales forecasting.
The production model must keep pace with market changes, so model accuracy is critically important.
Even though you have not changed the model itself since deploying it to production, the model’s accuracy has been gradually declining.
What is the most likely issue causing this continuous decline in accuracy?
Show answer
Correct answer: B. Lack of model retraining
Model retraining is the process of updating a model with new data to maintain its performance and accuracy.When accuracy gradually declines even though the model has not changed, the cause is model drift—the distribution of data such as the market shifting over time—and the remedy is periodic retraining.
Option A (data quality), Option C (too few layers), and Option D (incorrect split ratios) are factors that would produce a certain amount of error from the moment of deployment, and they do not explain the “continuous degradation” over time.
Because retraining is essential to keep up with environmental change, Option B is correct.
Overview of Vertex AI Model Monitoring
Q8Google Professional Machine Learning Engineer
Q8. You have developed an application that chains together multiple scikit-learn models to predict the optimal price for your company’s products.
The workflow logic is shown in the diagram below.
Team members also use each model individually in other solutions’ workflows.
You want to deploy this workflow while ensuring version control for both the individual models and the workflow as a whole.
The application must be able to scale down to zero.
You want to minimize compute resource usage and the manual effort of managing this setup.
What should you do?
Show answer
Correct answer: C. Expose each model as an endpoint in Vertex AI Endpoints. Use Cloud Run to integrate the workflow.
To deploy a multi-model workflow while balancing version control and low cost, it is efficient to expose each model with Vertex AI Endpoints and implement the integration layer with a lightweight serverless option.Exposing each model individually with Vertex AI Endpoints makes version control and reuse easy, and integrating the workflow with Cloud Run, which can scale to zero, minimizes resource usage when idle.
Option A requires building the integration layer yourself in a custom container, which carries high operational overhead, and Options B and D load model files directly, undermining per-model version control and reusability.
Overview of Cloud Run
Q9Google Professional Machine Learning Engineer
Q9. You work for a company that captures live footage of the checkout area in a retail store.
Using this live footage, you need to build a model that detects, in near real time, the number of customers waiting for service.
You want to implement the solution in as little time and with as little effort as possible.
How should you build the model?
Show answer
Correct answer: A. Use the Occupancy Analytics model in Vertex AI Vision.
The Occupancy Analytics model in Vertex AI Vision is a pretrained model designed to count people and vehicles within video frames.It provides standard features such as counting within an active zone, line-crossing counts, and dwell detection, so it can meet the need to detect the number of waiting customers in near real time with minimal effort.
Options C and D require your own annotation and training, which is labor-intensive, and Option B is a separate model specialized in person/vehicle detection and is not as directly suited as occupancy analytics.
Note that Vertex AI Vision has now been renamed Agent Platform Vision, but this model and the correct answer remain valid today.
Agent Platform Vision: Occupancy Analytics Tutorial
Q10Google Professional Machine Learning Engineer
Q10. You are training a custom language model for your company using a large dataset.
You plan to use the Reduction Server strategy on Vertex AI.
You need to configure the worker pools for the distributed training job.
What should you do?
Show answer
Correct answer: B. Give the machines in the first two worker pools GPUs and use a container image that runs the training code. For the third worker pool, use the reductionserver container image with no accelerators and choose a machine type that emphasizes bandwidth.
Reduction Server is a fast all-reduce mechanism for GPUs that aggregates gradients from workers using a dedicated group of reducers.Because the reducers are meant to run on lightweight CPU VMs that are cheaper than GPUs, the correct configuration is to give the third worker pool no accelerators and choose a machine type that emphasizes bandwidth.
Reduction Server does not support TPUs, so configurations that use TPUs, as in Options C and D, are inappropriate.
Option A is wrong because it assigns unnecessary GPUs to the reducer pool, increasing cost.
Option B is correct: GPUs for the two worker pools, and CPU plus high bandwidth for the reducer pool.
Vertex AI: Distributed Training (Reduction Server)
