Latest [Oct 14, 2021] Professional-Machine-Learning-Engineer Exam with Accurate Google Professional Machine Learning Engineer PDF Questions
Take a Leap Forward in Your Career by Earning Google 72 Questions
NEW QUESTION 43
A Machine Learning Specialist is given a structured dataset on the shopping habits of a company's customer base. The dataset contains thousands of columns of data and hundreds of numerical columns for each customer. The Specialist wants to identify whether there are natural groupings for these columns across all customers and visualize the results as quickly as possible.
What approach should the Specialist take to accomplish these tasks?
- A. Run k-means using the Euclidean distance measure for different values of k and create an elbow plot.
- B. Embed the numerical features using the t-distributed stochastic neighbor embedding (t-SNE) algorithm and create a line graph.
- C. Embed the numerical features using the t-distributed stochastic neighbor embedding (t-SNE) algorithm and create a scatter plot.
- D. Run k-means using the Euclidean distance measure for different values of k and create box plots for each numerical column within each cluster.
Answer: A
NEW QUESTION 44
You need to build classification workflows over several structured datasets currently stored in BigQuery. Because you will be performing the classification several times, you want to complete the following steps without writing code: exploratory data analysis, feature selection, model building, training, and hyperparameter tuning and serving. What should you do?
- A. Use Al Platform to run the classification model job configured for hyperparameter tuning
- B. Use Al Platform Notebooks to run the classification model with pandas library
- C. Configure AutoML Tables to perform the classification task
- D. Run a BigQuery ML task to perform logistic regression for the classification
Answer: B
NEW QUESTION 45
You work for a global footwear retailer and need to predict when an item will be out of stock based on historical inventory dat a. Customer behavior is highly dynamic since footwear demand is influenced by many different factors. You want to serve models that are trained on all available data, but track your performance on specific subsets of data before pushing to production. What is the most streamlined and reliable way to perform this validation?
- A. Use the last relevant week of data as a validation set to ensure that your model is performing accurately on current data
- B. Use k-fold cross-validation as a validation strategy to ensure that your model is ready for production.
- C. Use the entire dataset and treat the area under the receiver operating characteristics curve (AUC ROC) as the main metric.
- D. Use the TFX ModelValidator tools to specify performance metrics for production readiness
Answer: D
NEW QUESTION 46
A data scientist needs to identify fraudulent user accounts for a company's ecommerce platform. The company wants the ability to determine if a newly created account is associated with a previously known fraudulent user.
The data scientist is using AWS Glue to cleanse the company's application logs during ingestion.
Which strategy will allow the data scientist to identify fraudulent accounts?
- A. Create a FindMatches machine learning transform in AWS Glue.
- B. Execute the built-in FindDuplicates Amazon Athena query.
- C. Create an AWS Glue crawler to infer duplicate accounts in the source data.
- D. Search for duplicate accounts in the AWS Glue Data Catalog.
Answer: A
Explanation:
Explanation/Reference: https://docs.aws.amazon.com/glue/latest/dg/machine-learning.html
NEW QUESTION 47
You are developing a Kubeflow pipeline on Google Kubernetes Engine. The first step in the pipeline is to issue a query against BigQuery. You plan to use the results of that query as the input to the next step in your pipeline. You want to achieve this in the easiest way possible. What should you do?
- A. Write a Python script that uses the BigQuery API to execute queries against BigQuery Execute this script as the first step in your Kubeflow pipeline
- B. Use the BigQuery console to execute your query and then save the query results Into a new BigQuery table.
- C. Locate the Kubeflow Pipelines repository on GitHub Find the BigQuery Query Component, copy that component's URL, and use it to load the component into your pipeline. Use the component to execute queries against BigQuery
- D. Use the Kubeflow Pipelines domain-specific language to create a custom component that uses the Python BigQuery client library to execute queries
Answer: B
NEW QUESTION 48
A company ingests machine learning (ML) data from web advertising clicks into an Amazon S3 data lake. Click data is added to an Amazon Kinesis data stream by using the Kinesis Producer Library (KPL). The data is loaded into the S3 data lake from the data stream by using an Amazon Kinesis Data Firehose delivery stream.
As the data volume increases, an ML specialist notices that the rate of data ingested into Amazon S3 is relatively constant. There also is an increasing backlog of data for Kinesis Data Streams and Kinesis Data Firehose to ingest.
Which next step is MOST likely to improve the data ingestion rate into Amazon S3?
- A. Decrease the retention period for the data stream.
- B. Increase the number of shards for the data stream.
- C. Increase the number of S3 prefixes for the delivery stream to write to.
- D. Add more consumers using the Kinesis Client Library (KCL).
Answer: B
Explanation:
Explanation/Reference:
NEW QUESTION 49
You are going to train a DNN regression model with Keras APIs using this code:
How many trainable weights does your model have? (The arithmetic below is correct.)
- A. 500*256*0 25+256*128*0 25+128*2 = 40448
- B. 501*256+257*128+2 = 161154
- C. 500*256+256*128+128*2 = 161024
- D. 501*256+257*128+128*2=161408
Answer: A
NEW QUESTION 50
You developed an ML model with Al Platform, and you want to move it to production. You serve a few thousand queries per second and are experiencing latency issues. Incoming requests are served by a load balancer that distributes them across multiple Kubeflow CPU-only pods running on Google Kubernetes Engine (GKE). Your goal is to improve the serving latency without changing the underlying infrastructure. What should you do?
- A. Significantly increase the max_batch_size TensorFlow Serving parameter
- B. Switch to the tensorflow-model-server-universal version of TensorFlow Serving
- C. Significantly increase the max_enqueued_batches TensorFlow Serving parameter
- D. Recompile TensorFlow Serving using the source to support CPU-specific optimizations Instruct GKE to choose an appropriate baseline minimum CPU platform for serving nodes
Answer: A
NEW QUESTION 51
You need to train a computer vision model that predicts the type of government ID present in a given image using a GPU-powered virtual machine on Compute Engine. You use the following parameters:
* Optimizer: SGD
* Image shape = 224x224
* Batch size = 64
* Epochs = 10
* Verbose = 2
During training you encounter the following error: ResourceExhaustedError: out of Memory (oom) when allocating tensor. What should you do?
- A. Reduce the image shape
- B. Change the learning rate
- C. Reduce the batch size
- D. Change the optimizer
Answer: C
NEW QUESTION 52
A company is using Amazon Polly to translate plaintext documents to speech for automated company announcements. However, company acronyms are being mispronounced in the current documents.
How should a Machine Learning Specialist address this issue for future documents?
- A. Output speech marks to guide in pronunciation.
- B. Create an appropriate pronunciation lexicon.
- C. Use Amazon Lex to preprocess the text files for pronunciation
- D. Convert current documents to SSML with pronunciation tags.
Answer: D
Explanation:
Explanation/Reference: https://docs.aws.amazon.com/polly/latest/dg/ssml.html
NEW QUESTION 53
You have been asked to develop an input pipeline for an ML training model that processes images from disparate sources at a low latency. You discover that your input data does not fit in memory. How should you create a dataset following Google-recommended best practices?
- A. Convert the images Into TFRecords, store the images in Cloud Storage, and then use the tf. data API to read the images for training
- B. Convert the images to tf .Tensor Objects, and then run tf. data. Dataset. from_tensors ().
- C. Create a tf.data.Dataset.prefetch transformation
- D. Convert the images to tf .Tensor Objects, and then run Dataset. from_tensor_slices{).
Answer: D
NEW QUESTION 54
A Data Scientist is working on an application that performs sentiment analysis. The validation accuracy is poor, and the Data Scientist thinks that the cause may be a rich vocabulary and a low average frequency of words in the dataset.
Which tool should be used to improve the validation accuracy?
- A. Scikit-leam term frequency-inverse document frequency (TF-IDF) vectorizer
- B. Natural Language Toolkit (NLTK) stemming and stop word removal
- C. Amazon Comprehend syntax analysis and entity detection
- D. Amazon SageMaker BlazingText cbowmode
Answer: A
Explanation:
Explanation/Reference: https://monkeylearn.com/sentiment-analysis/
NEW QUESTION 55
You are building a model to predict daily temperatures. You split the data randomly and then transformed the training and test datasets. Temperature data for model training is uploaded hourly. During testing, your model performed with 97% accuracy; however, after deploying to production, the model's accuracy dropped to 66%. How can you make your production model more accurate?
- A. Normalize the data for the training, and test datasets as two separate steps.
- B. Apply data transformations before splitting, and cross-validate to make sure that the transformations are applied to both the training and test sets.
- C. Split the training and test data based on time rather than a random split to avoid leakage
- D. Add more data to your test set to ensure that you have a fair distribution and sample for testing
Answer: B
NEW QUESTION 56
This graph shows the training and validation loss against the epochs for a neural network.
The network being trained is as follows:
* Two dense layers, one output neuron
* 100 neurons in each layer
* 100 epochs
* Random initialization of weights
Which technique can be used to improve model performance in terms of accuracy in the validation set?
- A. Increasing the number of epochs
- B. Early stopping
- C. Adding another layer with the 100 neurons
- D. Random initialization of weights with appropriate seed
Answer: A
NEW QUESTION 57
A Machine Learning Specialist is working with a large cybersecurity company that manages security events in real time for companies around the world. The cybersecurity company wants to design a solution that will allow it to use machine learning to score malicious events as anomalies on the data as it is being ingested. The company also wants be able to save the results in its data lake for later processing and analysis.
What is the MOST efficient way to accomplish these tasks?
- A. Ingest the data into Apache Spark Streaming using Amazon EMR, and use Spark MLlib with k-means to perform anomaly detection. Then store the results in an Apache Hadoop Distributed File System (HDFS) using Amazon EMR with a replication factor of three as the data lake.
- B. Ingest the data using Amazon Kinesis Data Firehose, and use Amazon Kinesis Data Analytics Random Cut Forest (RCF) for anomaly detection. Then use Kinesis Data Firehose to stream the results to Amazon S3.
- C. Ingest the data and store it in Amazon S3. Use AWS Batch along with the AWS Deep Learning AMIs to train a k-means model using TensorFlow on the data in Amazon S3.
- D. Ingest the data and store it in Amazon S3. Have an AWS Glue job that is triggered on demand transform the new data. Then use the built-in Random Cut Forest (RCF) model within Amazon SageMaker to detect anomalies in the data.
Answer: A
NEW QUESTION 58
A Machine Learning Specialist has completed a proof of concept for a company using a small data sample, and now the Specialist is ready to implement an end-to-end solution in AWS using Amazon SageMaker. The historical training data is stored in Amazon RDS.
Which approach should the Specialist use for training a model using that data?
- A. Move the data to Amazon DynamoDB and set up a connection to DynamoDB within the notebook to pull data in.
- B. Move the data to Amazon ElastiCache using AWS DMS and set up a connection within the notebook to pull data in for fast access.
- C. Write a direct connection to the SQL database within the notebook and pull data in
- D. Push the data from Microsoft SQL Server to Amazon S3 using an AWS Data Pipeline and provide the S3 location within the notebook.
Answer: D
NEW QUESTION 59
A Mobile Network Operator is building an analytics platform to analyze and optimize a company's operations using Amazon Athena and Amazon S3.
The source systems send data in .CSV format in real time. The Data Engineering team wants to transform the data to the Apache Parquet format before storing it on Amazon S3.
Which solution takes the LEAST effort to implement?
- A. Ingest .CSV data using Apache Kafka Streams on Amazon EC2 instances and use Kafka Connect S3 to serialize data as Parquet
- B. Ingest .CSV data using Apache Spark Structured Streaming in an Amazon EMR cluster and use Apache Spark to convert data into Parquet.
- C. Ingest .CSV data from Amazon Kinesis Data Streams and use Amazon Kinesis Data Firehose to convert data into Parquet.
- D. Ingest .CSV data from Amazon Kinesis Data Streams and use Amazon Glue to convert data into Parquet.
Answer: D
Explanation:
Explanation/Reference:
NEW QUESTION 60
You need to train a computer vision model that predicts the type of government ID present in a given image using a GPU-powered virtual machine on Compute Engine. You use the following parameters:
* Optimizer: SGD
* Image shape = 224x224
* Batch size = 64
* Epochs = 10
* Verbose = 2
During training you encounter the following error: ResourceExhaustedError: out of Memory (oom) when allocating tensor. What should you do?
- A. Reduce the image shape
- B. Change the learning rate
- C. Change the optimizer
- D. Reduce the batch size
Answer: C
NEW QUESTION 61
You are building a model to predict daily temperatures. You split the data randomly and then transformed the training and test datasets. Temperature data for model training is uploaded hourly. During testing, your model performed with 97% accuracy; however, after deploying to production, the model's accuracy dropped to 66%. How can you make your production model more accurate?
- A. Apply data transformations before splitting, and cross-validate to make sure that the transformations are applied to both the training and test sets.
- B. Normalize the data for the training, and test datasets as two separate steps.
- C. Add more data to your test set to ensure that you have a fair distribution and sample for testing
- D. Split the training and test data based on time rather than a random split to avoid leakage
Answer: C
NEW QUESTION 62
......
Authentic Best resources for Professional-Machine-Learning-Engineer Online Practice Exam: https://www.torrentvce.com/Professional-Machine-Learning-Engineer-valid-vce-collection.html
Practice To Professional-Machine-Learning-Engineer - TorrentVCE Remarkable Practice On your Google Professional Machine Learning Engineer Exam: https://drive.google.com/open?id=15IJ2rdqk4vRKhs-B0rBCTxaiQzfX5iRv