> ## Documentation Index
> Fetch the complete documentation index at: https://developers.techwolf.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# How to run an integration

## Quick reference

<Steps>
  <Step title="Ensure the integration is set up">
    Follow
    [How to set up an integration](/integrations/datasource-integrations/data-integrator/how-to-set-up-integration)
    to configure and save your integration first.
  </Step>

  <Step title="Trigger a run">
    Either upload a file matching the primary prefix (automatic trigger, when
    enabled) or manually start a run from the integration details page.
    This is only supported for file-based integrations.
  </Step>

  <Step title="Optionally dry run first">
    Use the dry run option to preview how many entities will be created,
    updated, and deleted — without making any API calls.
  </Step>

  <Step title="Monitor progress">
    Watch the progress bar and estimated time to completion on the
    integration details page.
  </Step>

  <Step title="Review results">
    After the run completes, check the summary and failure overview to see
    what succeeded and what needs attention.
  </Step>
</Steps>

## Triggering a run

A **run** is a single execution of an integration: one attempt to reconcile new
input data with the SkillEngine API. There are two ways to trigger a run.

### Automatic triggers

When an integration is **enabled**, uploading a file whose name matches the
primary file prefix automatically triggers a run, or in the case of a connector,
when the connector is triggered on its schedule. This is the standard mode for
production: a source system uploads a file on a recurring schedule, and the Data
Integrator picks it up and processes it without manual intervention.

<Info>
  Only files matching the **primary** file prefix trigger a run. Uploading
  files for extra (joined) file configurations does not trigger the
  integration — the Data Integrator always uses the latest available version
  of those files when a run starts.
</Info>

### Manual triggers

To trigger a run manually:

1. Navigate to the integration details page
   (`admin-portal/data-integrator/integrations/<integration_id>`).
2. Click the **Run Integration** button.
3. A dialog appears asking you to select an input file. It pre-fills with the
   latest file matching the configured primary prefix.
4. Choose whether to perform a **dry run** (see below).
5. Click **Run** to start.

Manual triggers are useful for testing new integrations, reprocessing specific
files, or running integrations that are disabled for automatic triggers.

## Dry runs

A dry run executes the full integration pipeline, but **pauses right before
sending requests to the SkillEngine API**.

### What a dry run shows you

At the pause point, the Data Integrator reports how many entities will be:

* **Created** — new entities that exist in the desired state but not yet in the
  API.
* **Updated** — existing entities whose data has changed.
* **Deleted** — entities that should no longer exist in the API.

This gives you a clear picture of the impact before any changes are made.

### Continuing from a dry run

After reviewing the dry run results, you can choose to **continue** the run.
This proceeds with sending the computed create, update, and delete requests to
the SkillEngine API — no need to start over.

## What happens during a run

When a run is triggered, the Data Integrator executes the following steps:

<Steps>
  <Step title="Transform">
    The input file is processed through the configured integration pipeline
    (input node → operations → output node), producing a transformed output
    that matches the SkillEngine API's expected format.
  </Step>

  <Step title="Full Picture">
    The new data is reconciled with the data that has been sent in previous
    uploads or connector triggers, to get the full picture of what needs to
    be in the API.
  </Step>

  <Step title="Diff">
    The diff is computed between the desired picture (all data previously
    sent) and the current API state, to determine what needs to be created,
    updated, or deleted.
  </Step>

  <Step title="Load">
    The computed changes are sent to the SkillEngine API. Each entity is
    individually created, updated, or deleted. It is right before this step
    an integration pauses when dry-run is enabled or when more than 5% of
    entities will be deleted.
  </Step>
</Steps>

### Resilience and automatic retry

A key property of the Data Integrator is that **failed operations are
automatically retried on subsequent runs**. Recorded failures are persisted.
On the next run, the difference between the desired state and the actual state
will include any previously failed operations — and attempt them again.

This means that even when sending delta files (only changes since the last run),
entities that failed in previous runs are retried alongside new changes. Over
time, the API state converges with the desired state without manual
intervention.

There is no functionality in place to detect repeated failures of the same data.
It is up to the user to review failures every once in a while and resolve them.

<Accordion title="Example: Automatic retry across runs">
  **Run 1:** You load 1,000 Employees. 990 succeed, 10 fail due to a transient
  API error.

  **Run 2 (next day):** You upload a new file with 5 changed Employees. The Data
  Integrator computes the diff and finds:

  * 5 updates (from the new file)
  * 10 creates (the Employees that failed yesterday)

  All 15 are sent to the API. If the transient issue is resolved, the 10
  previously failed Employees now succeed alongside the 5 updates.
</Accordion>

## Monitoring progress

While a run is in progress, the integration details page shows:

* **A progress bar** indicating how far along the reconciliation phase is.
* **An estimated time to completion** based on current throughput.
* **Error reporting** happens only at the end of a run.

<Warning>
  The progress bar is **relative to the current run**, not the entire dataset.
  If the run only needs to process 10 changes out of a million entities, the
  progress bar tracks those 10. See the FAQ below for more details.
</Warning>

## Reviewing results

After a run completes, the integration details page shows:

* **A summary** of how many entities were successfully created, updated, and
  deleted.
* **A progress bar** indicating the success rate of the current run.
* **A failure overview** listing any entities that failed, along with the API
  response for each failure.

Failures are not necessarily cause for alarm — the FAQ below explains common
failure types and which ones require action.

## FAQ

<AccordionGroup>
  <Accordion title="Why does the progress bar look low even though my integration is mostly successful?">
    The progress bar shows the success rate **relative to the current run**, not the
    total dataset. If your integration manages a million entities in the API but
    only needed to process 10 changes today, and 4 of those 10 failed, the bar
    shows 60% — even though over 99.99% of your data is correctly loaded.

    This design helps you focus on what changed in the current run rather than
    masking failures behind a large denominator.
  </Accordion>

  <Accordion title="What do the different failure types mean?">
    Not all failures are equal. Here is a breakdown of common types:

    **Server errors (5xx) and connection errors**

    Transient infrastructure issues. These are **automatically retried** on
    the next run and should not persist across multiple consecutive runs. If they
    do, there may be an underlying infrastructure issue.

    **Bad request (400)**

    The data sent to the API is incorrectly formatted or violates a validation rule.
    This typically indicates a problem in the integration configuration (e.g., a
    date in the wrong format, a required field that is empty). Review the error
    message in the failure overview and adjust your integration configuration.

    **Not Found (404) on create**

    The parent entity does not exist in the API. For example, creating a Skill Event
    fails with 404 if the referenced Employee has not been loaded yet. These
    failures **automatically resolve** once the parent entity exists — the Data
    Integrator retries them on every subsequent run.

    **Not Found (404) on update**

    The entity does not exist in the API. This can happen when the entity was deleted
    outside of the Data Integrator, or after a state reset. The Data Integrator records
    the entity in its state and will issue an **update** instead of a create on the
    next run.

    **Not Found (404) on delete**

    Treated as a **success**. The entity is already gone — the desired outcome is
    achieved. This commonly happens with cascading deletes: when an Employee is
    deleted, their Skill Events are automatically removed by the API. If the Skill
    Event integration then tries to delete those same events, the 404 is expected.

    **Conflict (409) on create**

    The entity already exists in the API but was not in the Data Integrator's state.
    This can happen when data was loaded outside of the Data Integrator, or after a
    state reset. The Data Integrator records the entity in its state and will issue
    an **update** instead of a create on the next run.
  </Accordion>

  <Accordion title="Why do certain entities keep failing across multiple runs?">
    If the same entities fail run after run, the most common causes are:

    1. **Missing parent entity** — a sub-entity (e.g., a Skill Event) depends on a
       parent (e.g., an Employee) that does not exist in the API. This can happen
       when the parent entity's integration has not run yet, or the parent failed
       in its own integration. Ensure the parent integration is running and
       succeeding.
    2. **Invalid data** — the source data has a formatting or validation problem
       that retrying alone cannot fix. Check the failure details for the specific
       API error message and correct the issue in your integration configuration or
       source data.

    If entities are failing you think should not, contact TechWolf Support.
  </Accordion>

  <Accordion title="How does the initial load work when switching to Delta delivery?">
    A common pattern is to start with a **Full Dump** to load all existing data into
    the API, then switch to **Delta** delivery for all subsequent runs.

    1. Configure the integration with **Full Dump** delivery.
    2. Upload the complete dataset and run the integration.
    3. Once the initial load succeeds, edit the integration and change the delivery
       type to **Delta**.
    4. From now on, upload only incremental change files (with an `operation`
       column).

    The Data Integrator retains the full dataset from the initial load. Each Delta run
    applies changes on top of it, so it always represents the complete desired state.
  </Accordion>

  <Accordion title="Can I increase the concurrency of a run?">
    By default, runs execute with a concurrency of 2 (two parallel API requests).
    Increasing concurrency is not currently self-service. If you need higher
    throughput for large initial loads or time-sensitive migrations, contact TechWolf
    Support to discuss temporarily increasing concurrency for your integration.
  </Accordion>

  <Accordion title="What happens if I disable an integration?">
    Disabling an integration prevents automatic triggers — uploading a file matching
    the primary prefix will **not** start a run. You can still trigger runs manually
    from the integration details page.

    This is useful when you want to test an integration with dry runs before letting
    it process files automatically, or when you need to temporarily pause an
    integration while making configuration changes.
  </Accordion>
</AccordionGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.