Trang chủ Professional Data Engineer 20 / 343
Quay lại bộ đề
Question #20 Topic 1
An external customer provides you with a daily dump of data from their database. The data flows into Google Cloud Storage GCS as comma-separated values(CSV) files. You want to analyze this data in Google BigQuery, but the data could have rows that are formatted incorrectly or corrupted. How should you build this pipeline?
A
Use federated data sources, and check data in the SQL query.
B
Enable BigQuery monitoring in Google Stackdriver and create an alert.
C
Import the data into BigQuery using the gcloud CLI and set max_bad_records to 0.
D
Run a Google Cloud Dataflow batch pipeline to import the data into BigQuery, and push errors to another dead-letter table for analysis.