> Blog >

Snowflake Raises dbt Project Limits to 100,000 Files and Enhances Data Clean Rooms: Scaling Modern Data Collaboration

Snowflake Raises dbt Project Limits to 100,000 Files and Enhances Data Clean Rooms: Scaling Modern Data Collaboration

Fred
October 7, 2026

On August 27, 2026, Snowflake delivered two practical advances for teams that build and share data at scale. The maximum number of files allowed in a deployed dbt project object rose from 20,000 to 100,000. Concurrently, Snowflake Data Clean Rooms received API version 17.8 updates that improve ML job runtime handling, request-status accuracy, and overall performance.

Together these changes reduce friction for large transformation projects and make secure, privacy-preserving collaboration more reliable—especially when machine-learning workloads are involved. For data-engineering and privacy leaders, the message is clear: the platform continues to remove operational ceilings that previously forced workarounds or project fragmentation.

Higher dbt Project File Limits: From 20,000 to 100,000

A dbt project object on Snowflake can now contain up to 100,000 files. The limit counts every file in the project directory and its subdirectories, including:

  • Source models, tests, macros, and seeds
  • Generated folders such as target/
  • Installed packages in dbt_packages/
  • Log directories

Previously the ceiling sat at 20,000 files. Large enterprise projects—especially those with extensive package dependencies, generated documentation, or multi-domain models—could approach or exceed that threshold, forcing teams to split projects artificially or keep portions of the codebase outside Snowflake.

The five-fold increase gives most mature dbt codebases comfortable headroom while still encouraging good hygiene. Snowflake documentation continues to advise splitting into logical sub-projects when approaching the limit and invites customers with still-larger needs to contact their account representative.

Key technical points

  • Limit applies to deployed dbt project objects (the versioned objects used for scheduled execution).
  • Workspaces and Git-connected development flows remain subject to the same deployed-object ceiling once a version is published.
  • Concurrent execution of multiple EXECUTE DBT PROJECT commands against the same object is still unsupported; intra-project parallelism via dbt threads and selective –select runs remain the recommended patterns.

Data Clean Rooms Updates (API 17.8)

The August 27 Clean Rooms release focused on operational reliability and ML-friendly defaults:

  • ML Jobs image_tag default — When a code specification is registered without an explicit image_tag, the collaboration now resolves to the latest available runtime image at the time the code spec is added. Previously the default was pinned to the static tag 2.8.1. Existing code specs pick up the newer default only when re-added to a collaboration; teams that need a fixed version can still set image_tag explicitly.
  • Corrected requestor status — In VIEW_UPDATE_REQUESTS, a requestor’s own entry now correctly remains in REQUESTED status until the request is processed, rather than prematurely showing APPROVED.
  • General performance improvements and bug fixes, plus continued refinement of private-preview features.

These changes sit on top of earlier August improvements (external and Iceberg tables in same-region cross-cloud collaborations, preset tables in templates, connector clean-up) and continue the platform’s push toward production-grade, multi-party analytics and ML.

Why These Changes Matter

For large-scale dbt users
Higher file limits remove a common reason to keep transformation logic outside the governed Snowflake environment. Teams can keep more of their DAG, packages, and generated artifacts inside dbt Projects on Snowflake, benefiting from native scheduling, role-based access, Horizon lineage, and Workspace editing without constant project-splitting gymnastics.

For Clean Rooms and privacy teams
Automatic resolution to the latest ML runtime reduces the risk of collaborations running on stale images. Clearer request-status semantics improve the collaboration workflow for both providers and consumers. Together they make Clean Rooms a more trustworthy foundation for joint analytics and model training on sensitive data.

Comparisons to Previous Capabilities

dbt
The jump from 20k to 100k files is one of the most visible capacity improvements since dbt Projects on Snowflake became generally available. Earlier limits forced architectural compromises that the new ceiling largely eliminates for all but the very largest monorepos.

Clean Rooms
Earlier releases expanded data types that could be shared, improved template authoring, and refined the Collaboration API. The August 27 updates are more operational—runtime currency and status correctness—but they address real friction points that appear once collaborations move from proof-of-concept into repeated production use.

Use Cases

  • Enterprise data platforms — Central analytics engineering teams maintain large, multi-domain dbt projects that include extensive macros, tests, and package trees.
  • Regulated collaboration — Financial services, healthcare, and retail partners run privacy-preserving overlap analysis or model training inside Clean Rooms, now with more current ML runtimes by default.
  • Hybrid development — Teams develop in Workspaces or Git, deploy versioned dbt project objects, and schedule them with tasks, staying inside a single governed boundary even as projects grow.
  • Cross-cloud Clean Rooms — Collaborations that already leverage Cross-Cloud Auto-Fulfillment benefit from clearer request handling and up-to-date ML job images.

Competitive Context

Transformation platforms and clean-room solutions compete on both capability and operational ease. Raising internal project limits and tightening the collaboration experience reduces the incentive to keep large transformation codebases or sensitive joint workloads on external orchestration or separate clean-room vendors. Snowflake’s combination of native dbt execution and first-party Clean Rooms continues to differentiate it for organizations that want transformation and secure multi-party analytics inside one control plane.

Actionable Insights for Data Engineering and Privacy Leaders

  • Audit existing dbt projects approaching the old 20,000-file threshold and plan consolidation or cleaner package management now that 100,000 files are supported.
  • For very large monorepos, still consider logical sub-projects for ownership and scheduling clarity; contact Snowflake if you need limits beyond 100k.
  • Review Clean Rooms code specs that omit image_tag and decide whether to re-add them (to pick up the latest runtime) or pin an explicit tag for reproducibility.
  • Update internal collaboration runbooks to reflect the corrected REQUESTED status behavior.
  • Treat dbt project objects and Clean Rooms as complementary: use the former for internal transformation scale and the latter for controlled external data collaboration.
  • Monitor file counts in deployed projects and ML job runtime versions as part of regular platform health checks.

What This Signals for Data Collaboration Maturity in 2026

Platforms are moving past “it works for small and medium projects” toward “it works for the real size of enterprise codebases and real multi-party workflows.” Higher dbt limits and more reliable Clean Rooms defaults are incremental but meaningful steps in that direction. Organizations that keep transformation logic and collaborative analytics inside a governed, scalable environment will find it easier to apply consistent security, lineage, and cost controls as both internal pipelines and external partnerships grow.

Conclusion

Snowflake’s August 27, 2026 updates—raising the dbt project file limit to 100,000 and shipping Data Clean Rooms API 17.8 improvements—remove practical friction for two of the most important patterns in modern data work: large-scale transformation and secure multi-party collaboration. Data engineering teams gain room to keep sophisticated dbt projects fully inside Snowflake; privacy and partnership teams gain more current ML runtimes and clearer collaboration status semantics.

Taken together, the changes reinforce Snowflake’s position as a platform that can host both the heavy lifting of internal analytics engineering and the controlled sharing required for external data collaboration—at the scale 2026 enterprise workloads demand.