Skip to content
Data Mapping
Guides
Scoping Monitors with Path Rules

Scoping Data Store Monitors with Path Rules

Path rules are per-monitor regular expressions that tell a data store monitor which databases, schemas, tables and fields to detect, which to skip, and which partitioned or sharded tables to treat as one. A rule acts during detection: anything it skips is never created as a staged resource, so it is never classified and never appears in the Action Center.

Use path rules when a monitor's settings are too coarse. The databases and excluded_databases settings work on whole databases only. Path rules work at every level, down to individual columns and nested fields.

Path rules are configured through the Astralis API. The Admin UI doesn't display or edit them yet.

Patterns on this page are shown as they appear in JSON request bodies, with every regular expression backslash doubled. See Escape patterns in JSON.

Choose between path rules and monitor settings

GoalUse
Scan only some databases, or leave whole databases outThe monitor's databases or excluded_databases setting. An include rule does the same and also works below the database level.
Skip schemas, tables or columns by name or naming conventionAn exclude rule
Scan only one schema or a handful of tablesOne or more include rules
Classify one partition of a date-partitioned or sharded table instead of every copyA collapse rule
Skip every BigQuery date-sharded table (events_20260901) entirelyThe default BigQuery behavior, or the monitor's sharded_table_pattern setting. A collapse rule keeps one shard per family instead.

Path rules apply to every data store monitor: PostgreSQL, MySQL, Microsoft SQL Server, Amazon RDS, Google Cloud SQL, BigQuery, Snowflake, DynamoDB, Amazon S3, ScyllaDB, Salesforce and Microsoft Purview. Creating a rule on a website, identity provider or cloud infrastructure monitor returns 400 Bad Request.

Understand resource paths

A rule's pattern is matched against the resource path: the staged resource's URN without its leading monitor key. A monitor with the key warehouse that detects the column email stores it as warehouse.analytics.public.users.email, so the path is analytics.public.users.email. Each . separates one level from the next.

The number of levels depends on the data store:

Data storePath shapeExample path
PostgreSQL, Snowflake, SQL Server, Cloud SQL for PostgreSQLdatabase.schema.table.columnanalytics.public.users.email
Amazon RDS for PostgreSQLinstance>database.schema.table.column (the > is encoded)orders-db%3Eanalytics.public.users.email
MySQLdatabase.database.table.column (the database is also the schema)shop.shop.orders.customer_id
Amazon RDS for MySQLinstance.database.table.columnorders-db.shop.orders.customer_id
BigQueryproject.dataset.table.column[.nested_field]acme-prod.analytics.events_20260901.payload.user_id
Amazon S3bucket.object_key.fieldacme-exports.exports%2Fusers%2Ecsv.email
DynamoDB (single dataset)dynamodb_container_schema.table.attributedynamodb_container_schema.users.email
DynamoDB (dataset per table)table.table.attributeusers.users.email
Salesforcesalesforce.object.fieldsalesforce.Contact.Email

To see real paths for your monitor, list its results with GET /api/v1/plus/discovery-monitor/results?staged_resource_urn={{staged_resource_urn}} (see Monitor API) and drop the first segment of each urn.

Names with special characters are percent-encoded

A path contains names as they're stored in the URN, not as they appear in the data store. Letters, digits, _ and - are kept as-is. Every other character, including ., spaces, / and $, is percent-encoded. This keeps a dot in a path an unambiguous level separator.

Name in the data storeAppears in the path as
customer datacustomer%20data
first.namefirst%2Ename
ORDERS$2026ORDERS%242026
exports/users.csv (S3 object key)exports%2Fusers%2Ecsv
events_2026-09-01events_2026-09-01

Write patterns against the encoded form. To skip the column first.name, match first%2Ename. A pattern containing first\\.name or a literal space never matches it. The hex digits match in either case, so %2e works as well as %2E.

Matching ignores case

Patterns match without regard to case, so ^analytics\\.public\\. matches Snowflake's ANALYTICS.PUBLIC.USERS as well as PostgreSQL's analytics.public.users. To require exact case for part of a pattern, wrap it in (?-i:...). For example, ^(?-i:ANALYTICS\\.PUBLIC\\.users)$ matches the quoted lowercase Snowflake table users but not USERS.

Understand how rules are applied

Each rule has an effect:

  • exclude skips a matching resource and everything beneath it.
  • include keeps only matching resources, everything beneath them and the ancestors leading to them. Everything else is skipped.
  • collapse treats sibling resources the pattern matches as shards of one family, keeps one representative per family and skips the rest.

Detection applies rules in this order:

  1. The monitor's databases and excluded_databases settings choose which databases to scan.
  2. Exclude always wins. A resource that an exclude rule matches, or that sits under one, is skipped regardless of any include rule or the order the rules were created in.
  3. Includes narrow the scope. With no include rules, everything not excluded is in scope. With one or more include rules, a resource must match one, sit under one, or lead to one. Several include rules form a union.
  4. Collapse groups what remains. A collapse rule only sees resources that include and exclude rules allow.
  5. The monitor's sharded_table_pattern setting (and BigQuery's built-in default) skips matching tables, except a table a collapse rule keeps as a representative.

How include and exclude patterns match

Include and exclude patterns are Python regular expressions (opens in a new tab) searched in the path. They match anywhere unless anchored, and a match on a resource covers its whole subtree.

  • An exclude pattern staging matches every path that contains staging anywhere: a schema staging, but also a table staging_orders and a column staging_flag. Anchor patterns to the level you mean.
  • ^ anchors to the start of the path and $ to the end. ^analytics$ names the database analytics exactly. ^analytics also matches analytics_staging.
  • Write [^.]+ for "any one name". ^[^.]+\\.[^.]+\\.audit_log$ matches a table called audit_log in any database and schema, and nothing at the column level.
  • An include pattern must start with ^; an exclude pattern doesn't have to. Detection scans top-down and uses the anchor to decide which databases and schemas lead to an included path without listing all their tables. A pattern such as ^prod|staging passes this check, but its second branch is unanchored and keeps everything; write ^(prod|staging)\\. instead.

How collapse patterns match

A collapse pattern must start with ^ and must match the whole path of each shard, as if it also ended with $. That's different from include and exclude, which match anywhere in the path: a collapse pattern has to spell out every level from the start of the path, for example ^analytics\\.billing\\.payment_summary_\\d{4}_\\d{2}$. A pattern written like an exclude, such as payment_summary_\\d{4}_\\d{2}$, is rejected with 422.

  • Families. Shards are grouped by parent and by the pattern's capture groups. ^analytics\\.raw\\.([^.]+)_\\d{4}_\\d{2}$ puts orders_2026_08 and orders_2026_09 in one family and refunds_2026_08 and refunds_2026_09 in another. A pattern without a capture group puts everything it matches under one parent into a single family.
  • One level per shard. A capture group that spans a . doesn't name a shard, so the match is ignored. Write wildcards as [^.]+, not .+. A collapse rule written for tables doesn't also collapse the columns beneath them.
  • Case. Captured text keeps the path's case, so ORDERS_2026_08 and orders_2026_09 are separate families.
  • Representative. Detection keeps the shard furthest through review, judged by its columns: a promoted shard first, then a reviewed one, then a classified one, then any other shard that already has a staged resource. A shard you've reviewed is the family's source of truth, and the choice made on the first scan sticks, so the family is classified once. When several shards are equally far along, or none has a staged resource yet, the greatest name in natural order wins (t_10 sorts after t_9). For year-first date suffixes, that's the newest shard; for MM_DD_YYYY suffixes it isn't. A muted shard is kept only if no other shard can be, so muting the representative hands the family to another shard on the next scan.
  • Overlapping rules. When several collapse rules match the same path, the rule with the lowest position claims it. position is assigned when a rule is created, so this is the rule created first; updating a rule doesn't change its position. A rule doesn't apply beneath a shard it names itself, but a second rule can still collapse resources inside a representative.
  • Shards that aren't kept are skipped like excluded resources. On a monitor's first scan they're never staged or promoted, so the promoted dataset contains only the representative. For shards that were already detected, see Rules and resources that were already detected.

Escape patterns in JSON

Patterns travel in JSON request bodies, where a backslash must be doubled. The regular expression ^analytics\.public\. is written "^analytics\\.public\\." in JSON, and \d is written \\d. Every pattern on this page is shown in its JSON form.

Add a path rule

  1. Find the paths you want to target. List the monitor's results with GET /api/v1/plus/discovery-monitor/results and drop the leading monitor key from each urn. If the monitor hasn't run yet, build the path from the data store's path shape.

  2. Create the rule. Send the pattern, action and effect to the monitor's path-rules endpoint. action is always detection.

    POST /api/v1/plus/discovery-monitor/{monitor_config_key}/path-rules
    curl -X POST '{{FIDES_URL}}/api/v1/plus/discovery-monitor/warehouse/path-rules' \
      -H 'Content-Type: application/json' \
      -H 'Authorization: Bearer {{FIDES_ACCESS_TOKEN}}' \
      -d '{
        "pattern": "^analytics\\.public\\.",
        "action": "detection",
        "effect": "include",
        "description": "Only the public schema of the analytics database"
      }'

    In the above example:

    • {{FIDES_URL}} is the URL to your Astralis server
    • {{FIDES_ACCESS_TOKEN}} is an access token with the discovery_monitor:update scope
    • warehouse is the monitor's key

    A successful request returns 201 Created with the rule:

    201 Created
    {
      "id": "mpr_a004f809-6201-46c8-930b-6f1514a2c654",
      "monitor_config_key": "warehouse",
      "pattern": "^analytics\\.public\\.",
      "action": "detection",
      "effect": "include",
      "description": "Only the public schema of the analytics database",
      "created_at": "2026-09-29T18:07:16.511876Z",
      "updated_at": "2026-09-29T18:07:16.511876Z",
      "position": 1
    }
  3. Run the monitor. Rules take effect on the monitor's next scan. To scan now, call POST /api/v1/plus/discovery-monitor/{monitor_config_key}/execute.

  4. Verify the results. Browse the monitor's results in the Action Center or through the results API. Excluded resources aren't listed, included resources and their ancestors are, and each collapsed family shows one table. If an include rule matched nothing, the scan detects no tables and logs a warning (see Troubleshoot path rules).

For every endpoint, field and status code, see Path Rules in the Monitor API reference.

Path rule cookbook

Each recipe shows the request body for POST /api/v1/plus/discovery-monitor/{monitor_config_key}/path-rules. The examples use a PostgreSQL-style database.schema.table.column path unless noted.

Scan only one database

Keep the analytics database and everything in it; skip every other database.

{ "pattern": "^analytics$", "action": "detection", "effect": "include" }
PathResult
analytics, analytics.public.users.emailDetected
analytics_staging.public.usersSkipped: $ stops the match at the end of the name
marketing.public.leadsSkipped

The monitor's databases setting does the same thing and is visible in the Admin UI, so prefer it for database-level scoping. Without the $, ^analytics would also keep analytics_staging.

Skip staging, scratch and temporary schemas

Skip schemas called staging or scratch, and any schema starting with tmp_, in every database.

{ "pattern": "^[^.]+\\.(staging|scratch|tmp_[^.]*)$", "action": "detection", "effect": "exclude" }
PathResult
analytics.staging, analytics.staging.orders.id, analytics.tmp_jdoeSkipped, with everything beneath them
analytics.STAGING.TSkipped: matching ignores case
analytics.staging_v2Detected: $ requires the whole schema name to match
analytics.public.staging_orders, analytics.public.users.staging_flagDetected: the pattern only names the schema level

For data stores whose paths start at the schema (S3, DynamoDB, Salesforce), drop the database segment: "^(staging|scratch|tmp_[^.]*)$".

Skip a table by name everywhere

Skip every table called audit_log, in any database and schema.

{ "pattern": "^[^.]+\\.[^.]+\\.audit_log$", "action": "detection", "effect": "exclude" }
PathResult
analytics.public.audit_log, analytics.public.audit_log.id, marketing.crm.AUDIT_LOGSkipped
analytics.public.audit_log_archiveDetected
analytics.public.users.audit_log (a column)Detected

The shorter "\\.audit_log$" also works for tables, but it skips any column called audit_log too. A pattern without the leading dot, such as "audit_log$", also matches customer_audit_log.

Skip a column by name across all tables

Skip the top-level column raw_payload in every table, so it's never sampled or classified.

{ "pattern": "^[^.]+\\.[^.]+\\.[^.]+\\.raw_payload$", "action": "detection", "effect": "exclude" }
PathResult
analytics.public.events.raw_payloadSkipped
analytics.public.events.meta.raw_payload (nested field)Detected
analytics.public.events.raw_payload_v2Detected
analytics.public.raw_payload (a table)Detected

To also skip nested fields called raw_payload at any depth, but still not tables, use:

{ "pattern": "^(?:[^.]+\\.){3}(?:[^.]+\\.)*raw_payload$", "action": "detection", "effect": "exclude" }

Include one schema but exclude a table inside it

Scan only analytics.public, without its audit_log table. Create two rules; the order doesn't matter.

{ "pattern": "^analytics\\.public\\.", "action": "detection", "effect": "include" }
{ "pattern": "^analytics\\.public\\.audit_log$", "action": "detection", "effect": "exclude" }
PathResult
analytics, analytics.publicDetected: they lead to the included tables
analytics.public.users, analytics.public.users.emailDetected
analytics.public.audit_log, analytics.public.audit_log.idSkipped: exclude wins
analytics.sales.orders, analytics.public_archive.x, marketing.public.leadsSkipped

An include rule can't bring back a resource under an excluded one: with an exclude on ^analytics\\.public$, an include on ^analytics\\.public\\.users$ keeps nothing.

An include that names a column keeps that column's table and skips its other columns. ^analytics\\.public\\.users\\.email$ keeps analytics.public.users with only email. Keep this in mind for patterns like ^[^.]+\\.[^.]+\\.[^.]+\\.email$: every table is on the way to a possible email column, so every table is still created, even those without one.

Match Snowflake uppercase identifiers

Snowflake stores unquoted identifiers in uppercase. Write patterns in either case; they match both.

{ "pattern": "^analytics\\.public\\.", "action": "detection", "effect": "include" }

This keeps ANALYTICS.PUBLIC.USERS.EMAIL. When a schema contains both a quoted lowercase table "users" and an unquoted USERS, scope exact case to the part that needs it:

{ "pattern": "^(?-i:ANALYTICS\\.PUBLIC\\.users)$", "action": "detection", "effect": "exclude" }

This skips ANALYTICS.PUBLIC.users and keeps ANALYTICS.PUBLIC.USERS. Put the ^ outside the group: an include rule written (?-i:^...) is rejected because it doesn't start with ^.

Match names that contain dots, spaces or slashes

Match the percent-encoded form. To skip the column first.name wherever it appears:

{ "pattern": "\\.first%2Ename$", "action": "detection", "effect": "exclude" }

To skip the table customer data in analytics.public:

{ "pattern": "^analytics\\.public\\.customer%20data$", "action": "detection", "effect": "exclude" }

To skip every S3 object under the tmp/ prefix of the acme-exports bucket:

{ "pattern": "^acme-exports\\.tmp%2F", "action": "detection", "effect": "exclude" }
PathResult
analytics.public.customer%20data.first%2EnameSkipped by the first two rules
acme-exports.tmp%2F2026%2Freport%2EcsvSkipped by the third rule
acme-exports.exports%2Fusers%2EcsvDetected

Classify one partition of a date-partitioned table

An ETL job writes a new payment_summary_YYYY_MM table every month. Classify one of them instead of every month's copy.

{
  "pattern": "^analytics\\.billing\\.payment_summary_\\d{4}_\\d{2}$",
  "action": "detection",
  "effect": "collapse",
  "description": "Monthly payment_summary partitions"
}
PathResult
analytics.billing.payment_summary_2026_10Detected: the newest shard becomes the representative on the first scan
analytics.billing.payment_summary_2026_08, ..._2026_09Skipped
analytics.billing.payment_summaryDetected: not a shard
analytics.billing.payment_summary_2026_11 (created next month)Skipped: 2026_10 already has a staged resource, so it stays the representative

For daily tables with hyphenated dates, such as events_2026-09-01, the hyphens need no encoding:

{ "pattern": "^analytics\\.raw\\.events_\\d{4}-\\d{2}-\\d{2}$", "action": "detection", "effect": "collapse" }

Collapse several partition families in one schema

Use a capture group for the part of the name that identifies the family. One rule then covers every date-suffixed table in the schema, with one representative per base name.

{ "pattern": "^analytics\\.raw\\.([^.]+)_\\d{4}_\\d{2}$", "action": "detection", "effect": "collapse" }
PathsResult
analytics.raw.orders_2026_08, analytics.raw.orders_2026_09One family; orders_2026_09 is kept
analytics.raw.refunds_2026_08, analytics.raw.refunds_2026_09One family; refunds_2026_09 is kept
analytics.raw.order_items_2026_09A family of one; kept
analytics.raw.customersDetected: not a shard

Without the capture group, ^analytics\\.raw\\.[^.]+_\\d{4}_\\d{2}$ would put orders and refunds in a single family and keep only one of them. Make the suffix specific: _\\d+$ would also collapse distinct tables such as step_1 and step_2.

Keep one shard of BigQuery date-sharded tables

BigQuery monitors skip date-sharded tables such as events_20260901 by default: tables whose names match ^(.+?)_\\d{6,8}$ are not detected at all. The monitor's sharded_table_pattern setting (in datasource_params) replaces that default with your own pattern. To detect and classify one shard per family instead, add a collapse rule. A collapse rule's representative is kept even when it matches the sharded table pattern.

For one dataset:

{ "pattern": "^acme-prod\\.analytics\\.([^.]+)_\\d{8}$", "action": "detection", "effect": "collapse" }

For every dataset in every project the monitor scans, including six-digit _YYYYMM shards:

{ "pattern": "^[^.]+\\.[^.]+\\.([^.]+)_(?:\\d{6}|\\d{8})$", "action": "detection", "effect": "collapse" }
PathsResult
acme-prod.analytics.events_20260901, ..._20260902One family; events_20260902 is kept
acme-prod.analytics.sessions_20260901, ..._20260902One family; sessions_20260902 is kept
acme-prod.analytics.usersDetected: not a shard
acme-prod.analytics.events_20260902.payload.user_idDetected as a nested field of the kept shard

Shards that match the sharded table pattern but no collapse rule are still skipped.

Collapse per-tenant schemas

A database has one schema per tenant, tenant_001 through tenant_500, all with the same tables. Detect and classify one tenant's schema.

{ "pattern": "^analytics\\.tenant_\\d{3}$", "action": "detection", "effect": "collapse" }

On the first scan, analytics.tenant_500 is kept with all its tables, and the other 499 schemas are skipped with everything in them. A tenant added later is skipped too, because tenant_500 already has a staged resource. For unpadded names such as tenant_2 and tenant_100, ^analytics\\.tenant_\\d+$ works the same way; natural ordering keeps tenant_100.

This recipe works well when every tenant schema really has the same tables. Tables that exist in only some tenants are never detected unless they're in the representative. The newest tenant may also be nearly empty. To choose the representative yourself, run the first scan with a temporary exclude that leaves only that tenant, then delete the exclude:

{ "pattern": "^analytics\\.tenant_(?!001$)\\d{3}$", "action": "detection", "effect": "exclude" }

With this exclude in place, only tenant_001 is detected and staged. After you delete it, tenant_001 stays the representative because it already has a staged resource. If the tenants have already been scanned, review or promote the tenant you want instead: a reviewed shard always wins.

⚠️

Adding this rule to a monitor that has already scanned the tenants deletes the staged resources of every tenant schema that was never classified, with all the tables and fields beneath it. See Rules and resources that were already detected.

Partitioned tables inside the kept tenant can be collapsed with a second rule that spells out the tenant level, such as ^analytics\\.tenant_\\d{3}\\.([^.]+)_\\d{4}_\\d{2}$.

Collapse numbered column families

A wide table has columns attr_1 through attr_200 that hold the same kind of data. Classify one of them per table.

{ "pattern": "^[^.]+\\.[^.]+\\.[^.]+\\.attr_\\d+$", "action": "detection", "effect": "collapse" }
PathsResult
analytics.public.products.attr_1 … attr_200One family per table; attr_200 is kept
analytics.public.variants.attr_1, attr_2A separate family; attr_2 is kept
analytics.public.products.skuDetected: not a shard
analytics.public.attr_1 (a table)Detected: the pattern names the column level

The skipped columns aren't staged, so they're not in the promoted dataset and receive no data categories.

Change or remove a rule

Update a rule's pattern, effect or description with PATCH. Omitted fields are unchanged; "description": null clears the description.

PATCH /api/v1/plus/discovery-monitor/{monitor_config_key}/path-rules/{rule_id}
curl -X PATCH '{{FIDES_URL}}/api/v1/plus/discovery-monitor/warehouse/path-rules/mpr_a004f809-6201-46c8-930b-6f1514a2c654' \
  -H 'Content-Type: application/json' \
  -H 'Authorization: Bearer {{FIDES_ACCESS_TOKEN}}' \
  -d '{ "pattern": "^analytics\\.public\\.users$", "description": null }'

Delete a rule with DELETE /api/v1/plus/discovery-monitor/{monitor_config_key}/path-rules/{rule_id}. Resources it was skipping are detected as new additions on the next scan.

A pattern can be used once per monitor and action, whatever its effect: creating ^analytics\\.public\\. as an exclude when it already exists as an include returns 409 Conflict. Patterns that differ only in case, such as ^PROD\\. and ^prod\\., are separate rules that match the same paths.

Rules and resources that were already detected

Include and exclude rules only prevent resources from being created. When you add one to a monitor that has already run:

  • Staged resources that the rule now skips keep their status, classifications and place in the results tree. They're not rescanned or updated, and they're not reported as removed.
  • To take them out of review, mute them.
  • Resources the rule skips that weren't detected before are never created.

A collapse rule also cleans up shards that were already detected. On the next scan:

  • Detection picks the representative as described in How collapse patterns match: a shard you've promoted or reviewed is kept over newer ones.
  • Every other shard that nothing has happened to yet is deleted, with everything beneath it, so the duplicates leave the results. That's a shard whose own status and every status beneath it are still addition: it was never classified, reviewed, promoted or muted.
  • Any other shard, such as one that was classified, reviewed, promoted or muted, or has a classification in progress or an error, keeps its staged resource, status and classifications, so no classification or review work is lost. They're not updated by later scans, and promoted ones stay in the promoted dataset until you remove them. To take classified duplicates out of review, mute them.
  • Shards that appear later are skipped.

If you delete the collapse rule, the next scan detects the deleted shards again as new additions, and they're classified again.

Troubleshoot path rules

SymptomCause and fix
422: "An include rule's pattern must start with '^'..."Add ^ at the very start. For case-sensitive includes, write ^(?-i:...), not (?-i:^...).
422: "A collapse rule's pattern must start with '^' and match the whole path of each shard..."Rewrite the pattern from the start of the path, spelling out every level: ^analytics\\.raw\\.events_\\d{8}$ rather than events_\\d{8}. This also applies when you change an existing exclude rule's effect to collapse.
422: "Invalid regex pattern: nested quantifiers (e.g. (a+)+) are not allowed; found '...'"The message names the repeated group, such as (?:_\\d+)+. Spell out the repetition instead: _\\d{4}_\\d{2} for a date suffix, or _\\d+(?:_\\d+)?(?:_\\d+)? for up to three numeric parts.
422: "Invalid regular expression: ..."The pattern doesn't compile. Check that every \ is doubled in JSON and brackets are balanced.
422 for patterns over 500 charactersSplit the rule into several shorter rules.
409 ConflictThe monitor already has a rule with exactly this pattern. Update that rule's effect instead.
A scan finds no tables after adding an include ruleThe include matches nothing. The worker logs "has include path rules but no table matched them". Compare the pattern with real paths, including percent-encoded names.
A collapse rule has no effectThe pattern doesn't match the whole path of any shard. Check that it spells out every level, and use [^.]+ rather than .+ in capture groups: a capture that spans a . doesn't name a shard.
A scan fails with "took longer than 1.0s to match"The pattern backtracks too much. Simplify alternations inside repeated groups.

Each scan with path rules logs one summary line in the worker log, such as Monitor warehouse path rules skipped 1532 resources and kept 48 tables, and one Collapsing N shards into <urn> line per collapsed family.