The hardest decision in Microsoft Fabric is often not which interface to open. It is where a data product should live, who should transform it, and which storage and query surfaces should serve its consumers. A lakehouse may be the right foundation when Spark processing and open Delta tables are central to the workflow, while a warehouse can be a better home for teams that rely primarily on relational modeling and T-SQL development. For the Microsoft DP-600 exam, you need to reason about those choices and their consequences for analytic datasets, governance and semantic models.
Fabric’s OneLake provides a common storage foundation for different experiences, but “one lake” does not mean every team should place everything in one workspace or one set of raw tables. Analytics architecture still requires ownership boundaries, reliable transformations, discoverable business definitions and careful access control. A well-designed lakehouse is not simply a landing folder for extracted files; it is a deliberate path from source events to trusted, reusable information.
Choose the lakehouse for its processing model
A Fabric lakehouse brings together files, Delta tables, Spark-based engineering and a SQL analytics endpoint over relevant tables. Its major advantage is that different analytic workloads can operate around an open table format rather than requiring a separate copy for every compute engine. Data engineering teams can use notebooks and Spark transformations to prepare large or varied datasets; downstream analysts can query curated Delta tables through SQL and semantic models.
That architecture differs from a Fabric warehouse. The warehouse is a relational T-SQL-first experience suited to familiar SQL development patterns, stored procedures and warehouse-oriented serving. A lakehouse’s SQL analytics endpoint is primarily a read-oriented query surface for Delta data, not the place to perform arbitrary inserts and updates to lakehouse tables. When transformation logic requires writing lakehouse tables, use the supported engineering path rather than treating the SQL endpoint as a conventional transactional database.
The choice should follow team skills and workload requirements. A data science team processing raw telemetry with PySpark may naturally own a lakehouse. A finance reporting team that needs carefully governed dimensional tables and SQL-based development may prefer a warehouse serving layer. Hybrid designs are legitimate: lakehouse bronze and silver layers can feed a warehouse gold layer. Microsoft explicitly supports these combinations, so the task is to justify the flow instead of declaring one Fabric item the winner for every problem.
Make bronze, silver and gold do different jobs
The medallion pattern helps separate source fidelity from business interpretation. A bronze layer retains ingested records and source context so that unexpected changes can be investigated. Silver applies quality rules, type normalization, deduplication and enrichment. Gold represents stable business entities and measures intended for consumption. These labels are useful only when each layer has a real contract; three directories containing near-identical data merely add cost and confusion.
For an online retailer, bronze might contain raw order events, payment callbacks and fulfillment updates. Silver could resolve duplicate events, standardize timestamps and currencies, and attach stable product and customer keys. Gold would publish order-line facts and relevant dimensions at a known grain. If a payment adjustment arrives late, the engineering process must specify how it updates the trusted records. A report should not silently count both the initial order and its retry as two sales.
Physical separation is a design choice. Individual workspaces can provide stronger operational and permission boundaries, especially when different groups manage ingestion, transformation and reporting. Separate lakehouses per layer may clarify ownership, while a smaller environment may use a more compact arrangement if its controls remain understandable. Whatever the layout, document where bad records are quarantined, how reprocessing works and which transformations are authoritative.
Understand what Delta tables contribute
A Delta table stores data files and transactional metadata so readers and writers can coordinate consistent table state. The transaction log matters because a lakehouse is not merely a collection of Parquet files with a friendly name. Schema changes, incremental writes and maintenance operations have different implications when a table supports downstream queries. Teams need conventions for schema evolution, partitions, file maintenance and the lifecycle of obsolete data.
Table grain should be established early. If a customer-events table contains a mixture of session starts, purchases and refunds with different definitions of “amount,” downstream analysts will struggle to build reliable measures. Refine entities until every row represents an understood observation. Maintain key relationships, data type consistency and event time semantics. Preparing these structures upstream makes the eventual semantic model both smaller and more reliable.
Be careful with file layout. Excessively small Delta files can increase metadata and query overhead, while inappropriate partitioning creates many sparsely populated directories. Optimize based on real read and write patterns, not a fashionable partition key. Maintenance must be coordinated with consumers that rely on current Delta snapshots, and retention policies should account for audit or recovery requirements rather than indiscriminately deleting history.
Use OneLake shortcuts to avoid unnecessary copies
A shortcut makes data stored elsewhere appear under a OneLake path, allowing supported analytics engines to access it without a full local duplication. This can help a finance team refer to curated data owned by another workspace or to compatible external storage. It does not turn an untrusted source into a trusted one; ownership, freshness, authorization and schema contracts remain at the source. A broken shortcut or revoked permission may affect several consumers at once.
Consider a supply-chain lakehouse that needs a central product dimension curated by the commercial team. A shortcut can expose that table for analysis without maintaining a daily copy. The benefit is reduced duplication and consistent reference data. The risk is a tighter dependency: if the producer renames a column or changes permissions, consumers can fail immediately. Use contracts and dependency tracking before treating zero-copy access as operationally free.
Shortcuts also influence security design. Granting access to an item or workspace does not automatically mean every user should have direct access to all underlying files. Evaluate how OneLake data controls, SQL endpoint permissions and semantic-model permissions apply to the access route being used. Security must be tested from each actual consumption path, because controls on one query surface do not necessarily constrain every other engine reading the same storage.
Know when the SQL analytics endpoint helps
The lakehouse SQL analytics endpoint exposes eligible Delta tables through a familiar T-SQL interface. It is particularly useful for exploration, validation and reporting-oriented queries over data prepared by Spark. Analysts can use relational joins, aggregations and SQL views to express consuming needs without owning the data-engineering notebooks. The endpoint’s read-oriented nature is an architectural feature: it keeps primary table writes under a controlled engineering process.
Automatic table discovery does not make every file available as a relational object. Files in a general-purpose folder may not appear as SQL tables; supported Delta tables and their registration matter. When a dataset seems invisible from SQL, first check its format and location, then metadata synchronization and permissions. Re-running an ingestion job simply because the SQL surface has not discovered a table may introduce duplicates or delay the investigation.
The SQL endpoint is also a useful bridge between engineering and analytics. For example, a curated fact table may be produced in the lakehouse and then queried to validate row counts, uniqueness and totals before a semantic model is attached. The validation logic can become part of a quality gate so that incorrect datasets do not reach executive dashboards merely because a notebook completed without throwing an exception.
Decide where the star schema should be built
Semantic models work best when the meaning of their data is clear. Fact tables describe events at a defined grain; dimensions describe the entities used to filter and group them. Some transformations belong in the lakehouse or warehouse, where they can be tested and reused across many consumers. Others, such as business-specific measures and presentation semantics, belong in the semantic model. The trade-off is not simply performance. It is whether multiple teams agree on the definition.
If each report independently reconstructs net revenue from raw events, differences in refunds, timing and exchange rates can create contradictory numbers. Establish the cleaned order and payment facts upstream, then define the agreed analytical measure once for reuse. A semantic model can handle flexible calculations over trusted tables, but it should not become the only place where every source defect is repaired. This division between platform preparation and model semantics is also why DP-700 Fabric data engineering skills complement analytics engineering rather than replacing them.
The storage mode of the semantic model adds another architectural choice. Import provides an in-memory snapshot refreshed on a schedule. DirectQuery issues requests to a source engine under supported conditions. Direct Lake can work directly with Delta-backed data and load needed columns into memory. Each has different performance, freshness and security considerations. The correct choice follows query behavior and operational boundaries, not simply the fact that the data currently resides in Fabric.
Build freshness and recovery into the design
“Near real time” is not a meaningful guarantee until the team defines timestamps and tolerances. Source arrival, ingestion, transformation, Delta commit, semantic-model framing and report query all contribute to the age of an answer. A sales dashboard that promises fifteen-minute freshness should track each stage so a delayed payment feed is distinguishable from a model-refresh issue. A green pipeline run does not prove that the input contained the expected new records.
Reprocessing needs its own design. Idempotent loads, stable business keys, watermarks, change detection and explicit late-arriving-data rules reduce the chance that a replay doubles the numbers. If a source drops a column unexpectedly, the workflow should surface that change rather than quietly replace critical data with nulls. Test recovery from missed schedules, permission failures and partial writes. Explain the expected consistency between layers so users know which tables are preliminary and which are certified.
Turn architecture choices into DP-600 decisions
DP-600 scenarios often describe a stakeholder need rather than naming a Fabric component. Translate the requirement: Who owns the raw data? Is large-scale Spark transformation expected? Does the business need writable T-SQL development, a read-oriented query surface or a reusable semantic model? Can the table be referenced through a shortcut, and is zero-copy access allowed by security policy? Then ask what happens on the failure path and how the result will be monitored.
Keep the broader Microsoft data and Fabric certifications in view: analytics engineering is about creating trustworthy assets that other people can use, not merely knowing which menu creates a lakehouse. A sound design explains the movement of a record from source to report, the ownership of each transformation and the permission boundary at each access surface. That explanation is far more durable than memorizing an example workspace layout.