Möchten Sie dieMicrosoft DP-750 Zertifizierungsprüfung mühlos bestehen? Dann sind die Fragenkataloge zur Microsoft DP-750 Zertifizierung aus ZertSoft unerlässlich. Die Fragenpool zur Microsoft DP-750 Zertifizierungsprüfung aus ZertSoft werden von den erfahrenen Experten durch ständige Praxis entworfen, sie sind eine Kommbination aus Fragen und Antworten. Deswegen ist die Webseite ZertSoft die Beste. Wählen Sie ZertSoft, wartet eine schönere Zukunft auf Sie da.
| Section | Weight | Objectives |
|---|---|---|
| Set up and configure an Azure Databricks environment | 15-20% | - Create and configure Azure Databricks workspaces
|
| Secure and govern Unity Catalog objects | 15-20% | - Implement governance and security
|
| Prepare and process data | 30-35% | - Ingest and transform data
|
| Deploy and maintain data pipelines and workloads | 30-35% | - Manage production workloads
|
Wir ZertSoft bieten die besten Service an immer vom Standpunkt der Kunden aus. 24/7 online Kundendienst, kostenfreie Demo der Microsoft DP-750, vielfältige Versionen, einjährige kostenlose Aktualisierung der Microsoft DP-750 Prüfungssoftware sowie die volle Rückerstattung beim Durchfall usw. Das alles ist der Grund dafür, dass wir ZertSoft zuverlässig ist. Wenn Sie die Microsoft DP-750 Prüfung mit Hilfe unserer Produkte bestehen, hoffen wir Ihnen, unsere gemeisame Anstrengung nicht zu vergessen!
32. Frage
Note: This section contains one or more sets of questions with the same scenario and problem. Each question presents a unique solution to the problem. You must determine whether the solution meets the stated goals. More than one solution in the set might solve the problem. It is also possible that none of the solutions in the set solve the problem.
After you answer a question in this section, you will NOT be able to return. As a result, these questions do not appear on the Review Screen.
You have an Azure Databricks workspace that is enabled for Unity Catalog and contains a Delta table named Orders.
You load the Orders table into an Apache Spark DataFrame named df.
You need to create a DataFrame that excludes rows where the order amount is null.
Solution: You run the following expression.
df.filter(df.order_amount != None)
Does this meet the goal?
Antwort: A
Begründung:
Correct:
* You run the following expression.
df.dropna(subset=["order_amount"])
The expression df.dropna(subset=["order_amount"]) is an appropriate and effective way to exclude rows where order_amount is null.
* You run the following expression.
df.filter(df.order_amount.isNotNull())
To exclude rows where the order amount is null, you can use the isNotNull() method or a SQL expression within the filter() or where() functions.Here are the standard, appropriate expressions:
Option 1: Python/PySpark API (Recommended)
pythondf_clean = df.filter(df["order_amount"].isNotNull())
Incorrect:
* You run the following expression.
df.fillna(0, subset=['order_amount'])
* You run the following expression.
df.filter(df.order_amount != None)
Reference:
https://www.geeksforgeeks.org/python/filter-pyspark-dataframe-columns-with-none-or-null-values/
https://learn.microsoft.com/en-us/azure/databricks/pyspark/reference/classes/dataframe/dropna
33. Frage
You have an Azure Databricks workspace that is enabled for Unity Catalog.
You have a Lakeflow Spark Declarative Pipelines (SDP) pipeline that writes numerical data to a table named Table1 by using a data quality validation rule named rule1.
You need to modify rule1 to meet the following requirements:
- Ensure that amount is always greater than 0.
- Prevent an update to Table1 from being committed when data that
violates rule1 is detected.
Which statement should you execute?
Antwort: B
Begründung:
To ensure the data validation rule forces the pipeline update to abort and roll back transactions when data violates the condition, you must use a "fail" expectation operator. In Databricks Lakeflow Spark Declarative Pipelines (SDP), the command/syntax depends on whether your pipeline is written in Python or SQL.
Python Implementation
If your pipeline uses Python, apply the @dp.expect_or_fail decorator above your table definition (note: dp is the standard alias for the databricks.pipelines module in Lakeflow SDP):
dp.expect_or_fail("amount_greater_than_zero", "amount > 0")
Reference:
https://docs.databricks.com/aws/en/ldp/expectations
34. Frage
You have an Azure Databricks workspace that uses serverless compute.
You need to ingest data by using Lakeflow Jobs. New records must be processed as soon as they become available.
Which type of job trigger should you use for the ingestion?
Antwort: A
Begründung:
The best trigger type for this scenario is the Continuous trigger.
Immediate Processing: The Continuous trigger mode processes new records as soon as they arrive at the configured data sources. This matches your requirement to ingest and process records without waiting for an artificial time interval.
Native Serverless Support: When paired with serverless compute, Lakeflow Jobs efficiently manage resources by automatically scaling up or down according to the real-time stream volume.
Built-in Fault Tolerance: Continuous pipelines on Databricks automatically handle failures by retrying with an exponential backoff policy, keeping your automated ingestion operational 24/7 without manual intervention.
Reference:
https://docs.databricks.com/aws/en/jobs/continuous
35. Frage
A Delta table receives new fields in incoming JSON data. You want to automatically adapt without breaking pipelines. What should you enable?
Antwort: B
Begründung:
Auto Loader supports schema inference and schema evolution, allowing new columns to be added automatically during ingestion. This reduces pipeline maintenance and avoids failures due to schema drift. Disabling enforcement is unsafe and risks data corruption. Overwrite mode deletes existing data. Manual updates are not scalable.
36. Frage
You have an Azure Databricks workspace that is enabled for Unity Catalog and contains two managed Delta tables named sales.schema1.table1 and sales.schema1.table2.
sales.schema1.table1 contains sales data from the current year.
sales.schema1 .table2 contains historical data.
You need to load all the rows from sales.schema1.table1 into sales.schema1.table2. The solution must preserve any existing data in sales.schema1.table2 and minimize processing effort.
Which command should you run?
Antwort: C
Begründung:
To load all rows from one table into the other while preserving existing data and minimizing processing effort, you should use the SQL INSERT INTO statement.
Preserves Data: INSERT INTO appends new rows to the target table without modifying or deleting the existing data.
Lowest Processing Effort: It performs a direct data append at the storage level. Unlike MERGE INTO, it does not scan the target table for matches, saving significant compute time and costs.
Delta Lake Optimization: Because these are Delta tables, appending data simply writes new parquet files and commits them to the transaction log, making the operation fast and efficient.
Reference:
https://medium.com/@gema.correa/handling-schema-evolution-and-schema-compensation-in-databricks-lessons-from-the-field-7af8d915beef
37. Frage
......
Sorgen Sie noch um die Prüfungsunterlagen der Microsoft DP-750? Jetzt brauchen Sie keine Sorgen! Weil uns zu finden bedeutet, dass Sie schon die Schlüssel zur Prüfungszertifizierung der Microsoft DP-750 gefunden haben. Wir ZertSoft beschäftigen uns seit Jahren mit der Entwicklung der Software der IT-Zertifizierungsprüfung. Jetzt genießen wir einen guten Ruf weltweit. Wir bieten Ihnen die effektivsten Hilfe bei der Vorbereitung der Microsoft DP-750.
DP-750 Zertifizierungsfragen: https://www.zertsoft.com/DP-750-pruefungsfragen.html