[Q149-Q173] Try Professional-Data-Engineer Free Now! Real Exam Question Answers Updated [Mar 30, 2023]

Share

Try Professional-Data-Engineer Free Now! Real Exam Question Answers Updated [Mar 30, 2023]

Get Ready to Pass the Professional-Data-Engineer exam with Google Latest Practice Exam 

NEW QUESTION 149
You are training a spam classifier. You notice that you are overfitting the training data. Which three actions can you take to resolve this problem? (Choose three.)

  • A. Use a larger set of features
  • B. Increase the regularization parameters
  • C. Get more training examples
  • D. Reduce the number of training examples
  • E. Use a smaller set of features
  • F. Decrease the regularization parameters

Answer: A,C,F

Explanation:
Explanation/Reference:

 

NEW QUESTION 150
You want to use a database of information about tissue samples to classify future tissue samples as either normal or mutated. You are evaluating an unsupervised anomaly detection method for classifying the tissue samples. Which two characteristic support this method? (Choose two.)

  • A. You expect future mutations to have similar features to the mutated samples in the database.
  • B. You already have labels for which samples are mutated and which are normal in the database.
  • C. There are very few occurrences of mutations relative to normal samples.
  • D. You expect future mutations to have different features from the mutated samples in the database.
  • E. There are roughly equal occurrences of both normal and mutated samples in the database.

Answer: D,E

 

NEW QUESTION 151
Government regulations in your industry mandate that you have to maintain an auditable record of access to certain types of datA. Assuming that all expiring logs will be archived correctly, where should you store data that is subject to that mandate?

  • A. In a bucket on Cloud Storage that is accessible only by an AppEngine service that collects user information and logs the access before providing a link to the bucket.
  • B. In Cloud SQL, with separate database user names to each user. The Cloud SQL Admin activity logs will be used to provide the auditability.
  • C. Encrypted on Cloud Storage with user-supplied encryption keys. A separate decryption key will be given to each authorized user.
  • D. In a BigQuery dataset that is viewable only by authorized personnel, with the Data Access log used to provide the auditability.

Answer: D

 

NEW QUESTION 152
You want to archive data in Cloud Storage. Because some data is very sensitive, you want to use the "Trust No One" (TNO) approach to encrypt your data to prevent the cloud provider staff from decrypting your data. What should you do?

  • A. Specify customer-supplied encryption key (CSEK) in the .botoconfiguration file. Use gsutil cpto upload each archival file to the Cloud Storage bucket. Save the CSEK in a different project that only the security team can access.
  • B. Use gcloud kms keys create to create a symmetric key. Then use gcloud kms encryptto encrypt each archival file with the key. Use gsutil cpto upload each encrypted file to the Cloud Storage bucket.
    Manually destroy the key previously used for encryption, and rotate the key once.
  • C. Specify customer-supplied encryption key (CSEK) in the .botoconfiguration file. Use gsutil cpto upload each archival file to the Cloud Storage bucket. Save the CSEK in Cloud Memorystore as permanent storage of the secret.
  • D. Use gcloud kms keys createto create a symmetric key. Then use gcloud kms encryptto encrypt each archival file with the key and unique additional authenticated data (AAD). Use gsutil cp to upload each encrypted file to the Cloud Storage bucket, and keep the AAD outside of Google Cloud.

Answer: B

Explanation:
Explanation/Reference:

 

NEW QUESTION 153
You are designing an Apache Beam pipeline to enrich data from Cloud Pub/Sub with static reference data from BigQuery. The reference data is small enough to fit in memory on a single worker. The pipeline should write enriched results to BigQuery for analysis. Which job type and transforms should this pipeline use?

  • A. Batch job, PubSubIO, side-inputs
  • B. Streaming job, PubSubIO, BigQueryIO, side-inputs
  • C. Streaming job, PubSubIO, JdbcIO, side-outputs
  • D. Streaming job, PubSubIO, BigQueryIO, side-outputs

Answer: A

 

NEW QUESTION 154
You are developing a software application using Google's Dataflow SDK, and want to use conditional, for loops and other complex programming structures to create a branching pipeline. Which component will be used for the data processing operation?

  • A. Sink API
  • B. PCollection
  • C. Transform
  • D. Pipeline

Answer: C

Explanation:
In Google Cloud, the Dataflow SDK provides a transform component. It is responsible for the data processing operation. You can use conditional, for loops, and other complex programming structure to create a branching pipeline.

 

NEW QUESTION 155
What is the general recommendation when designing your row keys for a Cloud Bigtable schema?

  • A. Keep your row key as long as the field permits
  • B. Include multiple time series values within the row key
  • C. Keep the row keep as an 8 bit integer
  • D. Keep your row key reasonably short

Answer: D

Explanation:
A general guide is to, keep your row keys reasonably short. Long row keys take up additional memory and storage and increase the time it takes to get responses from the Cloud Bigtable server.
Reference: https://cloud.google.com/bigtable/docs/schema-design#row-keys

 

NEW QUESTION 156
Your neural network model is taking days to train. You want to increase the training speed. What can you do?

  • A. Increase the number of layers in your neural network.
  • B. Subsample your training dataset.
  • C. Increase the number of input features to your model.
  • D. Subsample your test dataset.

Answer: A

 

NEW QUESTION 157
Your company is in a highly regulated industry. One of your requirements is to ensure individual users have access only to the minimum amount of information required to do their jobs. You want to enforce this requirement with Google BigQuery. Which three approaches can you take? (Choose three.)

  • A. Restrict access to tables by role.
  • B. Use Google Stackdriver Audit Logging to determine policy violations.
  • C. Restrict BigQuery API access to approved users.
  • D. Segregate data across multiple tables or databases.
  • E. Disable writes to certain tables.
  • F. Ensure that the data is encrypted at all times.

Answer: A,B,C

Explanation:
Explanation/Reference:

 

NEW QUESTION 158
You are a retailer that wants to integrate your online sales capabilities with different in-home assistants, such as Google Home. You need to interpret customer voice commands and issue an order to the backend systems. Which solutions should you choose?

  • A. Dialogflow Enterprise Edition
  • B. Cloud Natural Language API
  • C. Cloud AutoML Natural Language
  • D. Cloud Speech-to-Text API

Answer: A

Explanation:
Dialogflow is used to do voice analytics on human computer interaction.

 

NEW QUESTION 159
You architect a system to analyze seismic data. Your extract, transform, and load (ETL) process runs as a series of MapReduce jobs on an Apache Hadoop cluster. The ETL process takes days to process a data set because some steps are computationally expensive. Then you discover that a sensor calibration step has been omitted. How should you change your ETL process to carry out sensor calibration systematically in the future?

  • A. Develop an algorithm through simulation to predict variance of data output from the last MapReduce job based on calibration factors, and apply the correction to all data.
  • B. Modify the transformMapReduce jobs to apply sensor calibration before they do anything else.
  • C. Introduce a new MapReduce job to apply sensor calibration to raw data, and ensure all other MapReduce jobs are chained after this.
  • D. Add sensor calibration data to the output of the ETL process, and document that all users need to apply sensor calibration themselves.

Answer: B

 

NEW QUESTION 160
You currently have a single on-premises Kafka cluster in a data center in the us-east region that is responsible for ingesting messages from IoT devices globally. Because large parts of globe have poor internet connectivity, messages sometimes batch at the edge, come in all at once, and cause a spike in load on your Kafka cluster.
This is becoming difficult to manage and prohibitively expensive. What is the Google-recommended cloud native architecture for this scenario?

  • A. An IoT gateway connected to Cloud Pub/Sub, with Cloud Dataflow to read and process the messages from Cloud Pub/Sub.
  • B. Cloud Dataflow connected to the Kafka cluster to scale the processing of incoming messages.
  • C. Edge TPUs as sensor devices for storing and transmitting the messages.
  • D. A Kafka cluster virtualized on Compute Engine in us-east with Cloud Load Balancing to connect to the devices around the world.

Answer: A

 

NEW QUESTION 161
The Dataflow SDKs have been recently transitioned into which Apache service?

  • A. Apache Beam
  • B. Apache Spark
  • C. Apache Kafka
  • D. Apache Hadoop

Answer: A

Explanation:
Dataflow SDKs are being transitioned to Apache Beam, as per the latest Google directive

 

NEW QUESTION 162
An organization maintains a Google BigQuery dataset that contains tables with user-level data. They want
to expose aggregates of this data to other Google Cloud projects, while still controlling access to the user-
level data. Additionally, they need to minimize their overall storage cost and ensure the analysis cost for
other projects is assigned to those projects. What should they do?

  • A. Create and share a new dataset and table that contains the aggregate results.
  • B. Create and share an authorized view that provides the aggregate results.
  • C. Create dataViewer Identity and Access Management (IAM) roles on the dataset to enable sharing.
  • D. Create and share a new dataset and view that provides the aggregate results.

Answer: C

Explanation:
Explanation/Reference:
Reference: https://cloud.google.com/bigquery/docs/access-control

 

NEW QUESTION 163
You are planning to migrate your current on-premises Apache Hadoop deployment to the cloud. You need to ensure that the deployment is as fault-tolerant and cost-effective as possible for long-running batch jobs. You want to use a managed service. What should you do?

  • A. Install Hadoop and Spark on a 10-node Compute Engine instance group with standard instances. Install the Cloud Storage connector, and store the data in Cloud Storage. Change references in scripts from hdfs:// to gs://
  • B. Deploy a Cloud Dataproc cluster. Use an SSD persistent disk and 50% preemptible workers. Store data in Cloud Storage, and change references in scripts from hdfs:// to gs://
  • C. Install Hadoop and Spark on a 10-node Compute Engine instance group with preemptible instances. Store data in HDFS. Change references in scripts from hdfs:// to gs://
  • D. Deploy a Cloud Dataproc cluster. Use a standard persistent disk and 50% preemptible workers. Store data in Cloud Storage, and change references in scripts from hdfs:// to gs://

Answer: D

 

NEW QUESTION 164
Case Study: 2 - MJTelco
Company Overview
MJTelco is a startup that plans to build networks in rapidly growing, underserved markets around the world. The company has patents for innovative optical communications hardware. Based on these patents, they can create many reliable, high-speed backbone links with inexpensive hardware.
Company Background
Founded by experienced telecom executives, MJTelco uses technologies originally developed to overcome communications challenges in space. Fundamental to their operation, they need to create a distributed data infrastructure that drives real-time analysis and incorporates machine learning to continuously optimize their topologies. Because their hardware is inexpensive, they plan to overdeploy the network allowing them to account for the impact of dynamic regional politics on location availability and cost. Their management and operations teams are situated all around the globe creating many-to- many relationship between data consumers and provides in their system. After careful consideration, they decided public cloud is the perfect environment to support their needs.
Solution Concept
MJTelco is running a successful proof-of-concept (PoC) project in its labs. They have two primary needs:
Scale and harden their PoC to support significantly more data flows generated when they ramp to more than 50,000 installations.
Refine their machine-learning cycles to verify and improve the dynamic models they use to control topology definition.
MJTelco will also use three separate operating environments ?development/test, staging, and production ?
to meet the needs of running experiments, deploying new features, and serving production customers.
Business Requirements
Scale up their production environment with minimal cost, instantiating resources when and where needed in an unpredictable, distributed telecom user community. Ensure security of their proprietary data to protect their leading-edge machine learning and analysis.
Provide reliable and timely access to data for analysis from distributed research workers Maintain isolated environments that support rapid iteration of their machine-learning models without affecting their customers.
Technical Requirements
Ensure secure and efficient transport and storage of telemetry data Rapidly scale instances to support between 10,000 and 100,000 data providers with multiple flows each.
Allow analysis and presentation against data tables tracking up to 2 years of data storing approximately
100m records/day
Support rapid iteration of monitoring infrastructure focused on awareness of data pipeline problems both in telemetry flows and in production learning cycles.
CEO Statement
Our business model relies on our patents, analytics and dynamic machine learning. Our inexpensive hardware is organized to be highly reliable, which gives us cost advantages. We need to quickly stabilize our large distributed data pipelines to meet our reliability and capacity commitments.
CTO Statement
Our public cloud services must operate as advertised. We need resources that scale and keep our data secure. We also need environments in which our data scientists can carefully study and quickly adapt our models. Because we rely on automation to process our data, we also need our development and test environments to work as we iterate.
CFO Statement
The project is too large for us to maintain the hardware and software required for the data and analysis.
Also, we cannot afford to staff an operations team to monitor so many data feeds, so we will rely on automation and infrastructure. Google Cloud's machine learning will allow our quantitative researchers to work on our high-value problems instead of problems with our data pipelines.
MJTelco needs you to create a schema in Google Bigtable that will allow for the historical analysis of the last 2 years of records. Each record that comes in is sent every 15 minutes, and contains a unique identifier of the device and a data record. The most common query is for all the data for a given device for a given day. Which schema should you use?

  • A. Rowkey: dateColumn data: device_id, data_point
  • B. Rowkey: date#data_pointColumn data: device_id
  • C. Rowkey: device_idColumn data: date, data_point
  • D. Rowkey: date#device_idColumn data: data_point
  • E. Rowkey: data_pointColumn data: device_id, date

Answer: E

 

NEW QUESTION 165
Your company produces 20,000 files every hour. Each data file is formatted as a comma separated values
(CSV) file that is less than 4 KB. All files must be ingested on Google Cloud Platform before they can be
processed. Your company site has a 200 ms latency to Google Cloud, and your Internet connection
bandwidth is limited as 50 Mbps. You currently deploy a secure FTP (SFTP) server on a virtual machine in
Google Compute Engine as the data ingestion point. A local SFTP client runs on a dedicated machine to
transmit the CSV files as is. The goal is to make reports with data from the previous day available to the
executives by 10:00 a.m. each day. This design is barely able to keep up with the current volume, even
though the bandwidth utilization is rather low.
You are told that due to seasonality, your company expects the number of files to double for the next three
months. Which two actions should you take? (Choose two.)

  • A. Assemble 1,000 files into a tape archive (TAR) file. Transmit the TAR files instead, and disassemble
    the CSV files in the cloud upon receiving them.
  • B. Redesign the data ingestion process to use gsutil tool to send the CSV files to a storage bucket in
    parallel.
  • C. Introduce data compression for each file to increase the rate file of file transfer.
  • D. Contact your internet service provider (ISP) to increase your maximum bandwidth to at least 100 Mbps.
  • E. Create an S3-compatible storage endpoint in your network, and use Google Cloud Storage Transfer
    Service to transfer on-premices data to the designated storage bucket.

Answer: B,E

 

NEW QUESTION 166
You are responsible for writing your company's ETL pipelines to run on an Apache Hadoop cluster. The pipeline will require some checkpointing and splitting pipelines. Which method should you use to write the pipelines?

  • A. Java using MapReduce
  • B. HiveQL using Hive
  • C. Python using MapReduce
  • D. PigLatin using Pig

Answer: C

 

NEW QUESTION 167
You designed a database for patient records as a pilot project to cover a few hundred patients in three clinics.
Your design used a single database table to represent all patients and their visits, and you used self-joins to generate reports. The server resource utilization was at 50%. Since then, the scope of the project has expanded.
The database must now store 100 times more patient records. You can no longer run the reports, because they either take too long or they encounter errors with insufficient compute resources. How should you adjust the database design?

  • A. Partition the table into smaller tables, with one for each clinic. Run queries against the smaller table pairs, and use unions for consolidated reports.
  • B. Add capacity (memory and disk space) to the database server by the order of 200.
  • C. Normalize the master patient-record table into the patient table and the visits table, and create other necessary tables to avoid self-join.
  • D. Shard the tables into smaller ones based on date ranges, and only generate reports with prespecified date ranges.

Answer: C

 

NEW QUESTION 168
You want to build a managed Hadoop system as your data lake. The data transformation process is composed of a series of Hadoop jobs executed in sequence. To accomplish the design of separating storage from compute, you decided to use the Cloud Storage connector to store all input data, output data, and intermediary data. However, you noticed that one Hadoop job runs very slowly with Cloud Dataproc, when compared with the on-premises bare-metal Hadoop environment (8-core nodes with 100-GB RAM).
Analysis shows that this particular Hadoop job is disk I/O intensive. You want to resolve the issue. What should you do?

  • A. Allocate more CPU cores of the virtual machine instances of the Hadoop cluster so that the networking bandwidth for each instance can scale up
  • B. Allocate additional network interface card (NIC), and configure link aggregation in the operating system to use the combined throughput when working with Cloud Storage
  • C. Allocate sufficient persistent disk space to the Hadoop cluster, and store the intermediate data of that particular Hadoop job on native HDFS
  • D. Allocate sufficient memory to the Hadoop cluster, so that the intermediary data of that particular Hadoop job can be held in memory

Answer: C

Explanation:
Its google recommended approach to use LocalDisk/HDFS to store Intermediate result and use Cloud Storage for initial and final results.

 

NEW QUESTION 169
You are developing a software application using Google's Dataflow SDK, and want to use conditional, for loops and other complex programming structures to create a branching pipeline. Which component will be used for the data processing operation?

  • A. Sink API
  • B. PCollection
  • C. Transform
  • D. Pipeline

Answer: C

Explanation:
Explanation
In Google Cloud, the Dataflow SDK provides a transform component. It is responsible for the data processing operation. You can use conditional, for loops, and other complex programming structure to create a branching pipeline.
Reference: https://cloud.google.com/dataflow/model/programming-model

 

NEW QUESTION 170
Your company has hired a new data scientist who wants to perform complicated analyses across very large datasets stored in Google Cloud Storage and in a Cassandra cluster on Google Compute Engine. The scientist primarily wants to create labelled data sets for machine learning projects, along with some visualization tasks.
She reports that her laptop is not powerful enough to perform her tasks and it is slowing her down. You want to help her perform her tasks. What should you do?

  • A. Host a visualization tool on a VM on Google Compute Engine.
  • B. Deploy Google Cloud Datalab to a virtual machine (VM) on Google Compute Engine.
  • C. Run a local version of Jupiter on the laptop.
  • D. Grant the user access to Google Cloud Shell.

Answer: D

 

NEW QUESTION 171
You're training a model to predict housing prices based on an available dataset with real estate properties.
Your plan is to train a fully connected neural net, and you've discovered that the dataset contains latitude and longitude of the property. Real estate professionals have told you that the location of the property is highly influential on price, so you'd like to engineer a feature that incorporates this physical dependency.
What should you do?

  • A. Provide latitude and longitude as input vectors to your neural net.
  • B. Create a feature cross of latitude and longitude, bucketize it at the minute level and use L2 regularization during optimization.
  • C. Create a feature cross of latitude and longitude, bucketize at the minute level and use L1 regularization during optimization.
  • D. Create a numeric column from a feature cross of latitude and longitude.

Answer: D

Explanation:
Explanation/Reference:
Reference https://cloud.google.com/bigquery/docs/gis-data

 

NEW QUESTION 172
When a Cloud Bigtable node fails, ____ is lost.

  • A. the last transaction
  • B. all data
  • C. the time dimension
  • D. no data

Answer: D

Explanation:
A Cloud Bigtable table is sharded into blocks of contiguous rows, called tablets, to help balance the workload of queries. Tablets are stored on Colossus, Google's file system, in SSTable format. Each tablet is associated with a specific Cloud Bigtable node. Data is never stored in Cloud Bigtable nodes themselves; each node has pointers to a set of tablets that are stored on Colossus. As a result:
Rebalancing tablets from one node to another is very fast, because the actual data is not copied. Cloud Bigtable simply updates the pointers for each node. Recovery from the failure of a Cloud Bigtable node is very fast, because only metadata needs to be migrated to the replacement node.
When a Cloud Bigtable node fails, no data is lost
Reference: https://cloud.google.com/bigtable/docs/overview

 

NEW QUESTION 173
......

Pass Your Next Professional-Data-Engineer Certification Exam Easily & Hassle Free: https://www.torrentvce.com/Professional-Data-Engineer-valid-vce-collection.html

Get Prepared for Your Professional-Data-Engineer Exam With Actual Google Study Guide!: https://drive.google.com/open?id=1607iE24v7f8Q_s61ikddW6fD3lgeYl76