TECHNOLOGY & CERTIFICATION EDITORIAL

Microsoft DP-750: Designing Data Models in Unity Catalog

In Azure Databricks, data modeling is inseparable from governance. A table design that makes a dashboard fast may create a privacy problem if the same data is exposed to the wrong role; a perfectly governed table may be unusable if its grain, history and update semantics are unclear. Microsoft’s DP-750 plan covers data engineering using Azure Databricks, including Unity Catalog modeling. The published Microsoft study-guide revision is scheduled for October 19, 2026, so its revised wording should not be presented as already effective on October 8. The exam target is not present in the approved Exam-Topics.info workbook; this draft therefore links to the relevant Microsoft certifications rather than inventing a DP-750 URL.

Model the business grain before the file format

An order-events source emits one row each time an order changes state, while finance expects one row per settled order. Treating those sources as interchangeable will inflate revenue and confuse reconciliation. Before building tables, declare what one row means, which keys identify a business entity, and whether updates represent corrections or new events. A raw landing zone can retain incoming facts with ingestion metadata, while curated tables apply documented business rules. The design should preserve a route back to original evidence when a stakeholder asks why a measure changed after a pipeline update.

Modeling decisions also affect consumers. Analysts may need a stable dimensional view, while operational applications expect an event stream or current-state entity table. One physical representation rarely serves all tasks optimally. Define the contract for each exposed dataset: key uniqueness, allowed null values, historical interpretation, update cadence and responsible owner. If business definitions change, version them or manage compatible transitions rather than silently altering a widely used table. Data engineering quality is not simply whether a SQL query returns rows; it is whether those rows mean what consumers believe they mean.

Unity Catalog organizes ownership and permission

Unity Catalog uses securable objects such as catalogs, schemas and tables to organize discovery, permissions and audit. Naming conventions should communicate domain and environment without embedding confidential information in object names. Determine which team owns the data, who may create or modify objects, who can read them, and which service principals operate pipelines. Use least-privilege grants appropriate to each job rather than a shared workspace-wide administrator identity. An analyst with SELECT access need not have permission to modify a table’s schema or change its access policies.

Lineage can support impact analysis by showing how downstream assets depend on upstream tables, but observed lineage is not a replacement for written business meaning. A diagram may show that a revenue table comes from a sales source; it does not explain whether canceled orders are excluded or when settlement corrections take effect. Combine technical lineage with data contracts and documented transformations. When reviewing permissions, test representative actual identities, including a user who should be denied a sensitive field. A permission design verified only while signed in as a workspace administrator is not a security test.

SCD choices express a history contract

Slowly changing dimensions offer different ways to retain changes in attributes such as customer segment or product category. Type 1 updates the current value without preserving every prior state in the dimension. Type 2 keeps records with effective periods or equivalent versioning so historical analyses can use the attribute that was valid at the time. Neither is universally correct. If a customer’s region is corrected because the original entry was wrong, finance may need a restatement. If the customer legitimately moved between regions, history may need to retain both locations. That distinction is a business rule that should be explicit before writing a merge statement.

A robust Type 2 process needs a stable business key, effective timestamps, a definition of current records, and careful treatment of late-arriving changes. Duplicate changes in the source stream can create overlapping history intervals if keys and ordering are weak. Test the model with a customer who changes twice on the same day, one arriving change that precedes already processed data, and a correction that should not create a new business event. Data modeling reviews should focus on those awkward cases because common happy-path examples rarely expose historical integrity problems.

Physical optimization should follow query behavior

Delta tables store transactional metadata and data files, enabling reliable updates and version-aware operations. Performance depends on file sizes, query predicates, clustering choices and workload shape. Too many small files can raise metadata and scheduling overhead, while an excessively wide denormalized table may force users to scan irrelevant columns. Liquid clustering or other supported layout strategies can improve particular access patterns, but the right choice should be derived from measured query plans and expected ingestion behavior. Avoid applying every optimization feature to every table by default.

Partitioning must be chosen carefully. A partition key with excessive cardinality can produce many small partitions and poor maintenance behavior. Conversely, a low-selectivity partition key may offer little benefit for a workload that filters primarily on another attribute. Test representative queries under realistic volume and record the physical-layout decision. Maintenance jobs such as file compaction and retention management also affect recovery and streaming consumers. An optimization that speeds up one report but breaks a downstream job is not an improvement to the data product as a whole.

Fine-grained access changes the query path

Unity Catalog can enforce row filters and column masks to restrict what particular users may see. In 2026, Databricks documents table-level policies as well as attribute-based access control approaches for broader, tag-driven enforcement. The choice affects maintainability: a manually configured rule on one table can become inconsistent when similar tables are added; an appropriately governed shared policy can apply controls more consistently. Verify how functions evaluate group membership, which workloads can access protected objects, and any performance or compatibility constraints of the chosen method.

Do not assume a secure view and a secured underlying table are equivalent. A curated view may be useful for presenting selected columns, whereas a catalog-level data-access policy can enforce a rule more centrally. Policies should be tested with different group combinations and denied-access cases. Audit logs can indicate who accessed data, but they do not replace prevention. When sharing data outside the organization, determine how provider and recipient permissions, table policies and the sharing mechanism interact. Data governance must be designed as part of the model, not attached after consumers have unrestricted access.

Handle schema evolution as a product change

An ingestion source adds a field, changes a type, or begins sending a new event subtype. Blind automatic schema evolution can propagate incompatible structures into curated tables. Define which changes are backward-compatible, which require a migration, and which should quarantine records until the source owner resolves ambiguity. Null defaults may be acceptable for a new optional descriptive field but dangerous for a key used to reconcile transactions. Track the version of source contracts and make it easy to compare schemas between development and production environments before deploying a pipeline.

Consumers need notice of meaningful semantic changes even when column names remain unchanged. A field called net_amount can silently change meaning if a new tax or discount rule is introduced. Maintain quality checks for expected ranges and referential relationships, and validate sample historical comparisons after each significant transformation change. An internal incident in which one metric diverges often originates in this gap between technical schema compatibility and business semantic stability.

A privacy review can change the physical design

A customer-data model might be optimized for broad analytics using a wide table that includes contact details, sensitive identifiers and order history. Even if only a handful of columns appear in dashboards, the underlying dataset may be available to users who should not see all fields. Rather than adding ad hoc masking at the end, the team can revisit the exposed grain and separate purposes: provide pseudonymous analytical keys where detailed identity is unnecessary, use governed tables for sensitive attributes, and publish a curated representation with documented access. This may simplify downstream security reviews and improve query behavior by reducing unnecessary data movement.

The model designer should test both business questions and denied-access cases. Can finance calculate regional revenue without accessing names? Can a support worker inspect an individual account only with authorized context? Does a data export preserve the same control intention? If access rules change, who validates downstream reports and sharing arrangements? The right Unity Catalog design is not the one with the largest number of security features enabled. It is the one whose structural choices make permissible access obvious, enforceable and auditable without undermining the data’s intended meaning.

A data-model design review that earns approval

For a regional sales dataset, require the team to explain each table’s grain, historical treatment, source ownership, permission model and common query pattern. Ask what happens after a duplicate event, late correction, customer-region move, access-policy change and schema addition. Compare the model with actual analytical questions, not just data source diagrams. Run performance and denied-access tests using representative users. Confirm that consumers can find the dataset, understand its business meaning and trace its origin without requesting privileged notebook access from the engineering team.

The practical DP-750 lesson is to treat Unity Catalog as part of the data-product architecture. A curated model is trustworthy when identity, history, physical layout and business semantics agree. The exam’s published October 19 changes should be checked again when that date arrives, but the design habit is durable: engineers must defend both what the data says and who is allowed to see it.

Back to Insights
Explore what matters. Knowledge that goes beyond the exam.
Explore ExamTopics