Scoping Data Store Monitors with Path Rules
Path rules are per-monitor regular expressions that tell a data store monitor which databases, schemas, tables and fields to detect, which to skip, and which partitioned or sharded tables to treat as one. A rule acts during detection: anything it skips is never created as a staged resource, so it is never classified and never appears in the Action Center.
Use path rules when a monitor's settings are too coarse. The databases and excluded_databases settings work on whole databases only. Path rules work at every level, down to individual columns and nested fields.
Path rules are configured through the Astralis API. The Admin UI doesn't display or edit them yet.
Patterns on this page are shown as they appear in JSON request bodies, with every regular expression backslash doubled. See Escape patterns in JSON.
Choose between path rules and monitor settings
| Goal | Use |
|---|---|
| Scan only some databases, or leave whole databases out | The monitor's databases or excluded_databases setting. An include rule does the same and also works below the database level. |
| Skip schemas, tables or columns by name or naming convention | An exclude rule |
| Scan only one schema or a handful of tables | One or more include rules |
| Classify one partition of a date-partitioned or sharded table instead of every copy | A collapse rule |
Skip every BigQuery date-sharded table (events_20260901) entirely | The default BigQuery behavior, or the monitor's sharded_table_pattern setting. A collapse rule keeps one shard per family instead. |
Path rules apply to every data store monitor: PostgreSQL, MySQL, Microsoft SQL Server, Amazon RDS, Google Cloud SQL, BigQuery, Snowflake, DynamoDB, Amazon S3, ScyllaDB, Salesforce and Microsoft Purview. Creating a rule on a website, identity provider or cloud infrastructure monitor returns 400 Bad Request.
Understand resource paths
A rule's pattern is matched against the resource path: the staged resource's URN without its leading monitor key. A monitor with the key warehouse that detects the column email stores it as warehouse.analytics.public.users.email, so the path is analytics.public.users.email. Each . separates one level from the next.
The number of levels depends on the data store:
| Data store | Path shape | Example path |
|---|---|---|
| PostgreSQL, Snowflake, SQL Server, Cloud SQL for PostgreSQL | database.schema.table.column | analytics.public.users.email |
| Amazon RDS for PostgreSQL | instance>database.schema.table.column (the > is encoded) | orders-db%3Eanalytics.public.users.email |
| MySQL | database.database.table.column (the database is also the schema) | shop.shop.orders.customer_id |
| Amazon RDS for MySQL | instance.database.table.column | orders-db.shop.orders.customer_id |
| BigQuery | project.dataset.table.column[.nested_field] | acme-prod.analytics.events_20260901.payload.user_id |
| Amazon S3 | bucket.object_key.field | acme-exports.exports%2Fusers%2Ecsv.email |
| DynamoDB (single dataset) | dynamodb_container_schema.table.attribute | dynamodb_container_schema.users.email |
| DynamoDB (dataset per table) | table.table.attribute | users.users.email |
| Salesforce | salesforce.object.field | salesforce.Contact.Email |
To see real paths for your monitor, list its results with GET /api/v1/plus/discovery-monitor/results?staged_resource_urn={{staged_resource_urn}} (see Monitor API) and drop the first segment of each urn.
Names with special characters are percent-encoded
A path contains names as they're stored in the URN, not as they appear in the data store. Letters, digits, _ and - are kept as-is. Every other character, including ., spaces, / and $, is percent-encoded. This keeps a dot in a path an unambiguous level separator.
| Name in the data store | Appears in the path as |
|---|---|
customer data | customer%20data |
first.name | first%2Ename |
ORDERS$2026 | ORDERS%242026 |
exports/users.csv (S3 object key) | exports%2Fusers%2Ecsv |
events_2026-09-01 | events_2026-09-01 |
Write patterns against the encoded form. To skip the column first.name, match first%2Ename. A pattern containing first\\.name or a literal space never matches it. The hex digits match in either case, so %2e works as well as %2E.
Matching ignores case
Patterns match without regard to case, so ^analytics\\.public\\. matches Snowflake's ANALYTICS.PUBLIC.USERS as well as PostgreSQL's analytics.public.users. To require exact case for part of a pattern, wrap it in (?-i:...). For example, ^(?-i:ANALYTICS\\.PUBLIC\\.users)$ matches the quoted lowercase Snowflake table users but not USERS.
Understand how rules are applied
Each rule has an effect:
excludeskips a matching resource and everything beneath it.includekeeps only matching resources, everything beneath them and the ancestors leading to them. Everything else is skipped.collapsetreats sibling resources the pattern matches as shards of one family, keeps one representative per family and skips the rest.
Detection applies rules in this order:
- The monitor's
databasesandexcluded_databasessettings choose which databases to scan. - Exclude always wins. A resource that an exclude rule matches, or that sits under one, is skipped regardless of any include rule or the order the rules were created in.
- Includes narrow the scope. With no include rules, everything not excluded is in scope. With one or more include rules, a resource must match one, sit under one, or lead to one. Several include rules form a union.
- Collapse groups what remains. A collapse rule only sees resources that include and exclude rules allow.
- The monitor's
sharded_table_patternsetting (and BigQuery's built-in default) skips matching tables, except a table a collapse rule keeps as a representative.
How include and exclude patterns match
Include and exclude patterns are Python regular expressions (opens in a new tab) searched in the path. They match anywhere unless anchored, and a match on a resource covers its whole subtree.
- An exclude pattern
stagingmatches every path that containsstaginganywhere: a schemastaging, but also a tablestaging_ordersand a columnstaging_flag. Anchor patterns to the level you mean. ^anchors to the start of the path and$to the end.^analytics$names the databaseanalyticsexactly.^analyticsalso matchesanalytics_staging.- Write
[^.]+for "any one name".^[^.]+\\.[^.]+\\.audit_log$matches a table calledaudit_login any database and schema, and nothing at the column level. - An include pattern must start with
^; an exclude pattern doesn't have to. Detection scans top-down and uses the anchor to decide which databases and schemas lead to an included path without listing all their tables. A pattern such as^prod|stagingpasses this check, but its second branch is unanchored and keeps everything; write^(prod|staging)\\.instead.
How collapse patterns match
A collapse pattern must start with ^ and must match the whole path of each shard, as if it also ended with $. That's different from include and exclude, which match anywhere in the path: a collapse pattern has to spell out every level from the start of the path, for example ^analytics\\.billing\\.payment_summary_\\d{4}_\\d{2}$. A pattern written like an exclude, such as payment_summary_\\d{4}_\\d{2}$, is rejected with 422.
- Families. Shards are grouped by parent and by the pattern's capture groups.
^analytics\\.raw\\.([^.]+)_\\d{4}_\\d{2}$putsorders_2026_08andorders_2026_09in one family andrefunds_2026_08andrefunds_2026_09in another. A pattern without a capture group puts everything it matches under one parent into a single family. - One level per shard. A capture group that spans a
.doesn't name a shard, so the match is ignored. Write wildcards as[^.]+, not.+. A collapse rule written for tables doesn't also collapse the columns beneath them. - Case. Captured text keeps the path's case, so
ORDERS_2026_08andorders_2026_09are separate families. - Representative. Detection keeps the shard furthest through review, judged by its columns: a promoted shard first, then a reviewed one, then a classified one, then any other shard that already has a staged resource. A shard you've reviewed is the family's source of truth, and the choice made on the first scan sticks, so the family is classified once. When several shards are equally far along, or none has a staged resource yet, the greatest name in natural order wins (
t_10sorts aftert_9). For year-first date suffixes, that's the newest shard; forMM_DD_YYYYsuffixes it isn't. A muted shard is kept only if no other shard can be, so muting the representative hands the family to another shard on the next scan. - Overlapping rules. When several collapse rules match the same path, the rule with the lowest
positionclaims it.positionis assigned when a rule is created, so this is the rule created first; updating a rule doesn't change its position. A rule doesn't apply beneath a shard it names itself, but a second rule can still collapse resources inside a representative. - Shards that aren't kept are skipped like excluded resources. On a monitor's first scan they're never staged or promoted, so the promoted dataset contains only the representative. For shards that were already detected, see Rules and resources that were already detected.
Escape patterns in JSON
Patterns travel in JSON request bodies, where a backslash must be doubled. The regular expression ^analytics\.public\. is written "^analytics\\.public\\." in JSON, and \d is written \\d. Every pattern on this page is shown in its JSON form.
Add a path rule
-
Find the paths you want to target. List the monitor's results with
GET /api/v1/plus/discovery-monitor/resultsand drop the leading monitor key from eachurn. If the monitor hasn't run yet, build the path from the data store's path shape. -
Create the rule. Send the pattern, action and effect to the monitor's
path-rulesendpoint.actionis alwaysdetection.POST /api/v1/plus/discovery-monitor/{monitor_config_key}/path-rulescurl -X POST '{{FIDES_URL}}/api/v1/plus/discovery-monitor/warehouse/path-rules' \ -H 'Content-Type: application/json' \ -H 'Authorization: Bearer {{FIDES_ACCESS_TOKEN}}' \ -d '{ "pattern": "^analytics\\.public\\.", "action": "detection", "effect": "include", "description": "Only the public schema of the analytics database" }'In the above example:
{{FIDES_URL}}is the URL to your Astralis server{{FIDES_ACCESS_TOKEN}}is an access token with thediscovery_monitor:updatescopewarehouseis the monitor's key
A successful request returns
201 Createdwith the rule:201 Created{ "id": "mpr_a004f809-6201-46c8-930b-6f1514a2c654", "monitor_config_key": "warehouse", "pattern": "^analytics\\.public\\.", "action": "detection", "effect": "include", "description": "Only the public schema of the analytics database", "created_at": "2026-09-29T18:07:16.511876Z", "updated_at": "2026-09-29T18:07:16.511876Z", "position": 1 } -
Run the monitor. Rules take effect on the monitor's next scan. To scan now, call
POST /api/v1/plus/discovery-monitor/{monitor_config_key}/execute. -
Verify the results. Browse the monitor's results in the Action Center or through the results API. Excluded resources aren't listed, included resources and their ancestors are, and each collapsed family shows one table. If an include rule matched nothing, the scan detects no tables and logs a warning (see Troubleshoot path rules).
For every endpoint, field and status code, see Path Rules in the Monitor API reference.
Path rule cookbook
Each recipe shows the request body for POST /api/v1/plus/discovery-monitor/{monitor_config_key}/path-rules. The examples use a PostgreSQL-style database.schema.table.column path unless noted.
Scan only one database
Keep the analytics database and everything in it; skip every other database.
{ "pattern": "^analytics$", "action": "detection", "effect": "include" }| Path | Result |
|---|---|
analytics, analytics.public.users.email | Detected |
analytics_staging.public.users | Skipped: $ stops the match at the end of the name |
marketing.public.leads | Skipped |
The monitor's databases setting does the same thing and is visible in the Admin UI, so prefer it for database-level scoping. Without the $, ^analytics would also keep analytics_staging.
Skip staging, scratch and temporary schemas
Skip schemas called staging or scratch, and any schema starting with tmp_, in every database.
{ "pattern": "^[^.]+\\.(staging|scratch|tmp_[^.]*)$", "action": "detection", "effect": "exclude" }| Path | Result |
|---|---|
analytics.staging, analytics.staging.orders.id, analytics.tmp_jdoe | Skipped, with everything beneath them |
analytics.STAGING.T | Skipped: matching ignores case |
analytics.staging_v2 | Detected: $ requires the whole schema name to match |
analytics.public.staging_orders, analytics.public.users.staging_flag | Detected: the pattern only names the schema level |
For data stores whose paths start at the schema (S3, DynamoDB, Salesforce), drop the database segment: "^(staging|scratch|tmp_[^.]*)$".
Skip a table by name everywhere
Skip every table called audit_log, in any database and schema.
{ "pattern": "^[^.]+\\.[^.]+\\.audit_log$", "action": "detection", "effect": "exclude" }| Path | Result |
|---|---|
analytics.public.audit_log, analytics.public.audit_log.id, marketing.crm.AUDIT_LOG | Skipped |
analytics.public.audit_log_archive | Detected |
analytics.public.users.audit_log (a column) | Detected |
The shorter "\\.audit_log$" also works for tables, but it skips any column called audit_log too. A pattern without the leading dot, such as "audit_log$", also matches customer_audit_log.
Skip a column by name across all tables
Skip the top-level column raw_payload in every table, so it's never sampled or classified.
{ "pattern": "^[^.]+\\.[^.]+\\.[^.]+\\.raw_payload$", "action": "detection", "effect": "exclude" }| Path | Result |
|---|---|
analytics.public.events.raw_payload | Skipped |
analytics.public.events.meta.raw_payload (nested field) | Detected |
analytics.public.events.raw_payload_v2 | Detected |
analytics.public.raw_payload (a table) | Detected |
To also skip nested fields called raw_payload at any depth, but still not tables, use:
{ "pattern": "^(?:[^.]+\\.){3}(?:[^.]+\\.)*raw_payload$", "action": "detection", "effect": "exclude" }Include one schema but exclude a table inside it
Scan only analytics.public, without its audit_log table. Create two rules; the order doesn't matter.
{ "pattern": "^analytics\\.public\\.", "action": "detection", "effect": "include" }{ "pattern": "^analytics\\.public\\.audit_log$", "action": "detection", "effect": "exclude" }| Path | Result |
|---|---|
analytics, analytics.public | Detected: they lead to the included tables |
analytics.public.users, analytics.public.users.email | Detected |
analytics.public.audit_log, analytics.public.audit_log.id | Skipped: exclude wins |
analytics.sales.orders, analytics.public_archive.x, marketing.public.leads | Skipped |
An include rule can't bring back a resource under an excluded one: with an exclude on ^analytics\\.public$, an include on ^analytics\\.public\\.users$ keeps nothing.
An include that names a column keeps that column's table and skips its other columns. ^analytics\\.public\\.users\\.email$ keeps analytics.public.users with only email. Keep this in mind for patterns like ^[^.]+\\.[^.]+\\.[^.]+\\.email$: every table is on the way to a possible email column, so every table is still created, even those without one.
Match Snowflake uppercase identifiers
Snowflake stores unquoted identifiers in uppercase. Write patterns in either case; they match both.
{ "pattern": "^analytics\\.public\\.", "action": "detection", "effect": "include" }This keeps ANALYTICS.PUBLIC.USERS.EMAIL. When a schema contains both a quoted lowercase table "users" and an unquoted USERS, scope exact case to the part that needs it:
{ "pattern": "^(?-i:ANALYTICS\\.PUBLIC\\.users)$", "action": "detection", "effect": "exclude" }This skips ANALYTICS.PUBLIC.users and keeps ANALYTICS.PUBLIC.USERS. Put the ^ outside the group: an include rule written (?-i:^...) is rejected because it doesn't start with ^.
Match names that contain dots, spaces or slashes
Match the percent-encoded form. To skip the column first.name wherever it appears:
{ "pattern": "\\.first%2Ename$", "action": "detection", "effect": "exclude" }To skip the table customer data in analytics.public:
{ "pattern": "^analytics\\.public\\.customer%20data$", "action": "detection", "effect": "exclude" }To skip every S3 object under the tmp/ prefix of the acme-exports bucket:
{ "pattern": "^acme-exports\\.tmp%2F", "action": "detection", "effect": "exclude" }| Path | Result |
|---|---|
analytics.public.customer%20data.first%2Ename | Skipped by the first two rules |
acme-exports.tmp%2F2026%2Freport%2Ecsv | Skipped by the third rule |
acme-exports.exports%2Fusers%2Ecsv | Detected |
Classify one partition of a date-partitioned table
An ETL job writes a new payment_summary_YYYY_MM table every month. Classify one of them instead of every month's copy.
{
"pattern": "^analytics\\.billing\\.payment_summary_\\d{4}_\\d{2}$",
"action": "detection",
"effect": "collapse",
"description": "Monthly payment_summary partitions"
}| Path | Result |
|---|---|
analytics.billing.payment_summary_2026_10 | Detected: the newest shard becomes the representative on the first scan |
analytics.billing.payment_summary_2026_08, ..._2026_09 | Skipped |
analytics.billing.payment_summary | Detected: not a shard |
analytics.billing.payment_summary_2026_11 (created next month) | Skipped: 2026_10 already has a staged resource, so it stays the representative |
For daily tables with hyphenated dates, such as events_2026-09-01, the hyphens need no encoding:
{ "pattern": "^analytics\\.raw\\.events_\\d{4}-\\d{2}-\\d{2}$", "action": "detection", "effect": "collapse" }Collapse several partition families in one schema
Use a capture group for the part of the name that identifies the family. One rule then covers every date-suffixed table in the schema, with one representative per base name.
{ "pattern": "^analytics\\.raw\\.([^.]+)_\\d{4}_\\d{2}$", "action": "detection", "effect": "collapse" }| Paths | Result |
|---|---|
analytics.raw.orders_2026_08, analytics.raw.orders_2026_09 | One family; orders_2026_09 is kept |
analytics.raw.refunds_2026_08, analytics.raw.refunds_2026_09 | One family; refunds_2026_09 is kept |
analytics.raw.order_items_2026_09 | A family of one; kept |
analytics.raw.customers | Detected: not a shard |
Without the capture group, ^analytics\\.raw\\.[^.]+_\\d{4}_\\d{2}$ would put orders and refunds in a single family and keep only one of them. Make the suffix specific: _\\d+$ would also collapse distinct tables such as step_1 and step_2.
Keep one shard of BigQuery date-sharded tables
BigQuery monitors skip date-sharded tables such as events_20260901 by default: tables whose names match ^(.+?)_\\d{6,8}$ are not detected at all. The monitor's sharded_table_pattern setting (in datasource_params) replaces that default with your own pattern. To detect and classify one shard per family instead, add a collapse rule. A collapse rule's representative is kept even when it matches the sharded table pattern.
For one dataset:
{ "pattern": "^acme-prod\\.analytics\\.([^.]+)_\\d{8}$", "action": "detection", "effect": "collapse" }For every dataset in every project the monitor scans, including six-digit _YYYYMM shards:
{ "pattern": "^[^.]+\\.[^.]+\\.([^.]+)_(?:\\d{6}|\\d{8})$", "action": "detection", "effect": "collapse" }| Paths | Result |
|---|---|
acme-prod.analytics.events_20260901, ..._20260902 | One family; events_20260902 is kept |
acme-prod.analytics.sessions_20260901, ..._20260902 | One family; sessions_20260902 is kept |
acme-prod.analytics.users | Detected: not a shard |
acme-prod.analytics.events_20260902.payload.user_id | Detected as a nested field of the kept shard |
Shards that match the sharded table pattern but no collapse rule are still skipped.
Collapse per-tenant schemas
A database has one schema per tenant, tenant_001 through tenant_500, all with the same tables. Detect and classify one tenant's schema.
{ "pattern": "^analytics\\.tenant_\\d{3}$", "action": "detection", "effect": "collapse" }On the first scan, analytics.tenant_500 is kept with all its tables, and the other 499 schemas are skipped with everything in them. A tenant added later is skipped too, because tenant_500 already has a staged resource. For unpadded names such as tenant_2 and tenant_100, ^analytics\\.tenant_\\d+$ works the same way; natural ordering keeps tenant_100.
This recipe works well when every tenant schema really has the same tables. Tables that exist in only some tenants are never detected unless they're in the representative. The newest tenant may also be nearly empty. To choose the representative yourself, run the first scan with a temporary exclude that leaves only that tenant, then delete the exclude:
{ "pattern": "^analytics\\.tenant_(?!001$)\\d{3}$", "action": "detection", "effect": "exclude" }With this exclude in place, only tenant_001 is detected and staged. After you delete it, tenant_001 stays the representative because it already has a staged resource. If the tenants have already been scanned, review or promote the tenant you want instead: a reviewed shard always wins.
Adding this rule to a monitor that has already scanned the tenants deletes the staged resources of every tenant schema that was never classified, with all the tables and fields beneath it. See Rules and resources that were already detected.
Partitioned tables inside the kept tenant can be collapsed with a second rule that spells out the tenant level, such as ^analytics\\.tenant_\\d{3}\\.([^.]+)_\\d{4}_\\d{2}$.
Collapse numbered column families
A wide table has columns attr_1 through attr_200 that hold the same kind of data. Classify one of them per table.
{ "pattern": "^[^.]+\\.[^.]+\\.[^.]+\\.attr_\\d+$", "action": "detection", "effect": "collapse" }| Paths | Result |
|---|---|
analytics.public.products.attr_1 … attr_200 | One family per table; attr_200 is kept |
analytics.public.variants.attr_1, attr_2 | A separate family; attr_2 is kept |
analytics.public.products.sku | Detected: not a shard |
analytics.public.attr_1 (a table) | Detected: the pattern names the column level |
The skipped columns aren't staged, so they're not in the promoted dataset and receive no data categories.
Change or remove a rule
Update a rule's pattern, effect or description with PATCH. Omitted fields are unchanged; "description": null clears the description.
curl -X PATCH '{{FIDES_URL}}/api/v1/plus/discovery-monitor/warehouse/path-rules/mpr_a004f809-6201-46c8-930b-6f1514a2c654' \
-H 'Content-Type: application/json' \
-H 'Authorization: Bearer {{FIDES_ACCESS_TOKEN}}' \
-d '{ "pattern": "^analytics\\.public\\.users$", "description": null }'Delete a rule with DELETE /api/v1/plus/discovery-monitor/{monitor_config_key}/path-rules/{rule_id}. Resources it was skipping are detected as new additions on the next scan.
A pattern can be used once per monitor and action, whatever its effect: creating ^analytics\\.public\\. as an exclude when it already exists as an include returns 409 Conflict. Patterns that differ only in case, such as ^PROD\\. and ^prod\\., are separate rules that match the same paths.
Rules and resources that were already detected
Include and exclude rules only prevent resources from being created. When you add one to a monitor that has already run:
- Staged resources that the rule now skips keep their status, classifications and place in the results tree. They're not rescanned or updated, and they're not reported as removed.
- To take them out of review, mute them.
- Resources the rule skips that weren't detected before are never created.
A collapse rule also cleans up shards that were already detected. On the next scan:
- Detection picks the representative as described in How collapse patterns match: a shard you've promoted or reviewed is kept over newer ones.
- Every other shard that nothing has happened to yet is deleted, with everything beneath it, so the duplicates leave the results. That's a shard whose own status and every status beneath it are still
addition: it was never classified, reviewed, promoted or muted. - Any other shard, such as one that was classified, reviewed, promoted or muted, or has a classification in progress or an error, keeps its staged resource, status and classifications, so no classification or review work is lost. They're not updated by later scans, and promoted ones stay in the promoted dataset until you remove them. To take classified duplicates out of review, mute them.
- Shards that appear later are skipped.
If you delete the collapse rule, the next scan detects the deleted shards again as new additions, and they're classified again.
Troubleshoot path rules
| Symptom | Cause and fix |
|---|---|
422: "An include rule's pattern must start with '^'..." | Add ^ at the very start. For case-sensitive includes, write ^(?-i:...), not (?-i:^...). |
422: "A collapse rule's pattern must start with '^' and match the whole path of each shard..." | Rewrite the pattern from the start of the path, spelling out every level: ^analytics\\.raw\\.events_\\d{8}$ rather than events_\\d{8}. This also applies when you change an existing exclude rule's effect to collapse. |
422: "Invalid regex pattern: nested quantifiers (e.g. (a+)+) are not allowed; found '...'" | The message names the repeated group, such as (?:_\\d+)+. Spell out the repetition instead: _\\d{4}_\\d{2} for a date suffix, or _\\d+(?:_\\d+)?(?:_\\d+)? for up to three numeric parts. |
422: "Invalid regular expression: ..." | The pattern doesn't compile. Check that every \ is doubled in JSON and brackets are balanced. |
422 for patterns over 500 characters | Split the rule into several shorter rules. |
409 Conflict | The monitor already has a rule with exactly this pattern. Update that rule's effect instead. |
| A scan finds no tables after adding an include rule | The include matches nothing. The worker logs "has include path rules but no table matched them". Compare the pattern with real paths, including percent-encoded names. |
| A collapse rule has no effect | The pattern doesn't match the whole path of any shard. Check that it spells out every level, and use [^.]+ rather than .+ in capture groups: a capture that spans a . doesn't name a shard. |
| A scan fails with "took longer than 1.0s to match" | The pattern backtracks too much. Simplify alternations inside repeated groups. |
Each scan with path rules logs one summary line in the worker log, such as Monitor warehouse path rules skipped 1532 resources and kept 48 tables, and one Collapsing N shards into <urn> line per collapsed family.