> For the complete documentation index, see [llms.txt](https://docs.veza.com/4yItIzMvkpAvMVFAamTf/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.veza.com/4yItIzMvkpAvMVFAamTf/integrations/configuration/extraction.md).

# Extraction and Discovery Intervals

Customize how often Veza updates your Access Graph

Veza periodically connects to integrated systems to maintain an up-to-date graph of authorization and relationships between entities in your environment. This process involves two main activities: discovering new data sources and extracting authorization metadata.

You can customize both intervals globally or per-integration to manage compute cost at the source system and to keep large data sources from delaying other integrations. For guidance on when to change an interval, see [General guidelines](#general-guidelines).

All integrations have a default `auto` setting. **Most integrations default to 1 hour**, including Workday, AWS, Azure, Okta, and other identity providers. Some integrations use extended intervals to optimize performance and reduce costs: SharePoint Sites (24h), Snowflake (6h), and Box (24h). Wiz also extracts every 24 hours, set on its own interval setting rather than as a built-in default.

The `auto` setting allows Veza to manage extraction timing based on system load and integration type. When you override with a manual setting, Veza will schedule extraction precisely at that interval regardless of system conditions.

{% hint style="info" %}
**Understanding intervals**: An interval is the minimum time between the **completion** of one run and the **start** of the next. It sets a floor rather than a fixed schedule, so actual timing also depends on system load, data source size, and queue position. A 1-hour extraction interval means the next run starts no sooner than 1 hour after the previous run finishes.
{% endhint %}

An administrator can change the default value for each integration, or set global extraction and discovery intervals.

## Discovery interval

* Determines how often Veza checks for newly added data sources in your integrated systems.
* The discovery interval can be set between 15 minutes and 30 days.

## Extraction interval

* Determines how often Veza collects authorization metadata to update entities in the Access Graph.
* Extraction intervals can be set between 1 hour and 30 days. Wiz is the single exception: its minimum is 24 hours.
* Gathers supported entities and their attributes, which can take some time for large data sources.
* More frequent extractions keep entity relationships and attributes closer to current, and increase load on the source system.

## How much each run collects

Most integrations perform a **full collection** on every run: each extraction gathers the complete set of in-scope objects. Two features change this, and they are often confused because both depend on audit logs. They do different things.

**Incremental extraction** changes *what* a run collects. Where an integration supports it, Veza reads the source's activity logs to identify what changed and applies those changes to a baseline established by an earlier full extraction. Veza uses incremental extraction for **Okta**, primarily to reduce API load on large tenants.

Incremental extraction does not change the interval: runs still start on the configured interval, and audit log collection can poll more frequently than the extraction itself. It is also conditional. Veza performs a full extraction when audit log collection is not enabled, when the audit log has fallen too far behind, when no baseline payload exists, or when the integration configuration changes. Veza also refreshes the baseline with a full extraction every seven days by default. When the audit log reports no changes, the run completes without calling the Okta API.

**Audit log scheduling** changes *when* a run happens, not what it collects. Rather than re-extracting on every interval, Veza collects audit logs and watches for management events, such as permission changes, group membership changes, and file and folder activity. When it sees one, it marks the data source out of date, and the next scheduled extraction performs a **full** update. Data sources with no detected changes are skipped. **SharePoint Online** uses this behavior.

Three details matter when planning around audit log scheduling:

* It applies to the *next* scheduled run, so the configured interval is still the floor. It is not an immediate trigger.
* File and folder activity, including file access and downloads, marks a data source out of date, so an active site is rarely skipped. Most sharing-link events do not.
* A full extraction runs at least once every seven days regardless of detected changes (or once per interval, when the interval is longer), and Veza reverts to interval-based extraction if audit log collection begins failing.

See [Enable Audit Logs for Okta](/4yItIzMvkpAvMVFAamTf/integrations/integrations/okta.md#enable-audit-logs-for-okta) and [Enable SharePoint integration](/4yItIzMvkpAvMVFAamTf/integrations/integrations/azure.md#3-enable-sharepoint-integration-optional).

## General guidelines

Leave intervals at the `auto` default unless a specific requirement calls for a change. Veza selects default intervals to stay within the API rate limits each vendor enforces, and a shorter interval can cause extractions to back off and take longer overall.

The central tradeoff is data freshness against cost and load on the source system:

* More frequent extraction keeps the Access Graph closer to current, and increases API calls and compute cost at the source.
* Less frequent extraction lowers cost and produces more stable performance, which suits high-volume sources such as SharePoint.

Configure a scheduled extraction when a downstream process depends on the timing, such as refreshing HR data before a Lifecycle Management workflow runs. Monitor the effect of any interval change on the integration **Details** page before applying the same change more widely.

The `auto` default suits most deployments. Snowflake and SharePoint warrant more consideration, because both carry longer default intervals for cost and scale reasons and the right interval depends on data volume, compute budget, and how current the data needs to be.

### Snowflake

Snowflake compute cost scales with query volume and frequency, so a shorter extraction interval raises cost. The default interval is 6 hours.

To reduce cost, lengthen the extraction interval, which lowers cost roughly in proportion to the reduction in frequency. You can also reduce the number of databases Veza queries with the **Database Allow List** and **Database Deny List** settings, described in [Limiting Extractions](/4yItIzMvkpAvMVFAamTf/integrations/configuration/limits.md).

Cost depends on how much data is queried and the compute required to return it, so assess the effect against your own environment rather than against a fixed estimate. See also [Adjust warehouse autosuspend to reduce uptime after extractions](/4yItIzMvkpAvMVFAamTf/integrations/integrations/snowflake.md#adjust-warehouse-autosuspend-to-reduce-uptime-after-extractions).

### SharePoint

The default extraction interval for SharePoint sites is 24 hours. In tenants with many sites and folders, a 7-day extraction interval paired with a 12-hour discovery interval can reduce total extraction time, if that freshness suits your organization.

Two settings reduce SharePoint extraction volume further:

* Clear **Gather personal sites** in the Azure integration's **Limit Services** settings, unless personal OneDrive sites are in scope.
* Enable audit log extraction, so that Veza connects to SharePoint Online for a full update only when the Office 365 Management Activity API reports a change. This requires enablement at the tenant level. If the option is not available in your tenant, contact your Veza support representative to request it. See [Enable SharePoint integration](/4yItIzMvkpAvMVFAamTf/integrations/integrations/azure.md#3-enable-sharepoint-integration-optional).

## If extraction is slow

When an integration takes longer to extract than the freshness you need, four levers are available, in rough order of effect:

* **Narrow the scope.** Several integrations limit how much each run collects. AWS can extract RDS metadata at the database level, skipping schemas and tables; Azure can skip personal SharePoint sites; SharePoint Server accepts site allow and deny lists. See [Limiting Extractions](/4yItIzMvkpAvMVFAamTf/integrations/configuration/limits.md).
* **Enable audit log extraction** where the source supports it. For Okta this enables incremental extraction, so each run collects only new and changed data. For SharePoint it lets Veza skip runs for data sources with no detected changes, rather than reducing what a run collects. Both reduce API load. See [How much each run collects](#how-much-each-run-collects).
* **Set a longer interval** for that integration. Base it on how often the source data actually changes rather than on a uniform value across integrations. See [How to Customize Intervals](#how-to-customize-intervals).
* **Move extraction to an off-peak window** so that it does not compete with business-hours traffic on the source system. For example, schedule a full Okta extraction for the weekend. See [Advanced Scheduling](#advanced-scheduling).

If an integration error persists after these changes, contact Veza support or your account team.

## Manual/on-demand extraction

In addition to scheduled extractions, you can manually trigger synchronization for individual data sources. This is particularly useful for immediate updates, testing, and time-sensitive changes, such as updating Access Graph metadata for new Access Reviews or Lifecycle Management workflows.

### How to trigger manual extraction

You can manually start extraction in two ways:

**From the All Data Sources page:**

1. Navigate to **Integrations** > **All Data Sources**
2. Filter to locate the data source you want to extract
3. Click the **Start Extraction** button in the Actions column
4. The extraction will be queued and the status will update to show progress

![Triggering a manual extraction from All Data Sources view.](https://1967633068-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F-MZDkWMxox3pekd0NsZJ%2Fuploads%2Fgit-blob-f55ded7be2da0059f81b477fd6b3f2787a7669a2%2Fdata-sources-start-extraction.png?alt=media)

**From Integration Details:**

1. Navigate to **Integrations** and select a specific integration
2. Click the **Data Sources** tab
3. Click the **Start Extraction** button next to the desired data source
4. Monitor the extraction status in the Status column

Manual extractions run with the same permissions and configuration as scheduled extractions. The next scheduled extraction will still occur at its regular interval.

{% hint style="info" %}
**Prioritization**: User-initiated extractions are automatically prioritized in the extraction queue with higher priority than scheduled extractions.

**Availability**: Manual extraction is available for all native integrations (e.g., Okta, AWS, Azure, Snowflake). After initiating extraction the status will show "Extraction In Progress" with a timestamp. You can monitor progress via the status indicators and click on error messages for full details.
{% endhint %}

### How to customize intervals

1. Open the **Administration** cog icon (at the bottom of the navigation sidebar).
2. Go to the **System Settings** tab.
3. Scroll down to the **Integrations** section.
4. Use the dropdown menus to set global values for discovery and extraction intervals.
   * Discovery Interval: How frequently to discover datasource instances.
   * Extraction Interval: How frequently to extract datasources.
5. To customize intervals for specific providers:
   * Use the *Search* field to filter for a specific integration.
   * Set custom values for individual providers as needed, from `auto` to any value between 1 hour and 30 days.

After updating an override, it will be listed beneath the global values, along with any other active overrides. Changes take effect immediately, with the next scheduled extraction/discovery following the new interval settings.

### Which setting applies

Veza resolves the interval for a data source in this order, and uses the first value that is set:

1. The override you set for that integration.
2. The integration's built-in default, where it has one.
3. The global value you set for all integrations.
4. The built-in global default of 1 hour.

A built-in default takes precedence over the global value, but only six data source types ship with one:

| Data source type                                           | Built-in default |
| ---------------------------------------------------------- | ---------------- |
| SharePoint                                                 | 24 hours         |
| Snowflake, Snowflake database, Snowflake volatile database | 6 hours          |
| Box enterprise, Box user                                   | 24 hours         |

For every other integration, changing the global extraction interval takes effect. To change the interval for one of the six above, set an override for that integration.

A per-data source extraction schedule, where one is configured, replaces interval scheduling entirely and takes precedence over all of the above. See [Advanced scheduling](#advanced-scheduling).

## Advanced scheduling

For data sources that need to extract at a specific time each day, Veza supports per-data source scheduling through a private API. A common reason to use it is to make sure HR system data is current before Lifecycle Management workflows run. There is no UI for this setting.

{% hint style="warning" %}
These endpoints are in the `private` namespace. They are not part of the public API contract and can change between releases.
{% endhint %}

Two capabilities are available:

* **Priority scheduling**: Queue a data source's extraction jobs ahead of lower-priority work. Use this when faster processing is needed but exact timing is not critical.
* **Scheduled extraction times**: Define exact times, and optionally days of the week, when extractions run.

### How a schedule interacts with the extraction interval

A data source with `scheduled_extraction_times` configured leaves interval scheduling entirely. Veza no longer consults that data source's extraction interval, and the schedule takes precedence over every value in [Which setting applies](#which-setting-applies).

Scheduled times are exact. Veza creates each scheduled job up to an hour before its configured time and holds it until that moment arrives, so a data source scheduled for `06:00:00` starts at 06:00 local time, plus the time a worker takes to pick the job up. This is the one case where extraction timing is a fixed wall-clock time rather than a delay measured from the end of the previous run.

Interval scheduling behaves differently. Veza evaluates interval eligibility on a recurring sweep, so an interval-based extraction can start a few minutes after its interval elapses.

If a scheduled time passes while an extraction for that data source is still running, Veza skips that run instead of queuing it. The data source extracts again at its next configured time. There is no catch-up run.

Discovery can be scheduled the same way. A `DISCOVERER` data source uses the same `scheduled_extraction_times` field, despite what the field name suggests.

### Constraints

* Times are specified in `HH:MM:SS` format; minutes must be `:00` or `:30`; times must be at least 1 hour apart
* A time zone must be provided in IANA format (for example, `America/New_York`)
* When `scheduled_extraction_times` is configured, `priority` must be `100`
* Day-of-week scheduling requires `scheduled_extraction_times`; use full uppercase day names (`MONDAY`–`SUNDAY`)
* By default, 100 data sources can hold scheduling configurations. Only Veza can raise this limit, so contact Veza support if you need more
* Supported for `EXTRACTOR` and `DISCOVERER` data source types only

{% hint style="info" %}
Check day names carefully. An unrecognized value in `scheduled_days_of_week` is dropped from the request instead of returning an error, so a misspelled day results in a schedule that omits that day.
{% endhint %}

### Configure scheduled extraction

```bash
curl -X POST "$BASE_URL/api/private/providers/datasources/{datasource_id}/scheduling_config" \
  -H "authorization: Bearer $VEZA_TOKEN" \
  -H "Content-Type: application/json" \
  --data-raw '{
    "priority": 100,
    "timezone": "America/New_York",
    "scheduled_extraction_times": ["06:00:00"],
    "scheduled_days_of_week": ["MONDAY", "TUESDAY", "WEDNESDAY", "THURSDAY", "FRIDAY"]
  }'
```

For priority only (no fixed time):

```bash
curl -X POST "$BASE_URL/api/private/providers/datasources/{datasource_id}/scheduling_config" \
  -H "authorization: Bearer $VEZA_TOKEN" \
  -H "Content-Type: application/json" \
  --data-raw '{"priority": 100}'
```

### View or remove configurations

```bash
# List all configured data sources
curl -X GET "$BASE_URL/api/private/providers/datasources/scheduling_configs" \
  -H "authorization: Bearer $VEZA_TOKEN"

# Remove a configuration (reverts to standard periodic scheduling)
curl -X DELETE "$BASE_URL/api/private/providers/datasources/{datasource_id}/scheduling_config" \
  -H "authorization: Bearer $VEZA_TOKEN"
```

The `datasource_id` is the UUID shown in the URL when viewing an integration's data sources, or returned by the [List Data Sources](/4yItIzMvkpAvMVFAamTf/developers/api/management/datasources/listdatasources.md) API. For full endpoint documentation and request/response schemas, see the [Scheduling Configuration API reference](/4yItIzMvkpAvMVFAamTf/developers/api/management/datasources/datasourceschedulingconfig.md).


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.veza.com/4yItIzMvkpAvMVFAamTf/integrations/configuration/extraction.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
