> ## Documentation Index
> Fetch the complete documentation index at: https://developers.techwolf.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Technical data flow

This page covers the parts shared by every file-based datasource specification
in this section: how files travel from the customer to TechWolf, where they land
on S3, and how to read the field tables on the
[Jobs](/integrations/datasource-integrations/non-standard-datasource-integrations/jobs),
[Employees](/integrations/datasource-integrations/non-standard-datasource-integrations/employees)
and
[Courses](/integrations/datasource-integrations/non-standard-datasource-integrations/courses)
pages.

For the broader picture of how file-based integrations work and how they compare
to API-based connectors, see
[Key concepts](/integrations/datasource-integrations/datasources-key-concepts).

## Technical data flow

The files described in the specification pages end up in the customer's TechWolf
**AWS S3** bucket, under the file-based ingest prefix. The file-based
integration reads them from there and syncs the data into the TechWolf
SkillEngine API. How the files reach S3 depends on the delivery method agreed
during implementation:

* **TechWolf-hosted SFTP server.** The customer (or their source system) pushes
  the files to the TechWolf SFTP server. Each SFTP user's root directory maps to
  a directory in TechWolf's S3 infrastructure, so an upload lands on S3
  directly. TechWolf provides the credentials. Server addresses and static IP
  addresses for allowlisting are listed under
  [SFTP: TechWolf-hosted SFTP server](/integrations/reference/sftp/sftp#techwolf-hosted-sftp-server).
* **Customer-hosted SFTP server.** The customer exports the files to their own
  SFTP server. TechWolf connects to it on a schedule, pulls the files, and
  writes them to S3 according to the folder mappings configured in the Console.
  Setup is described in
  [Self-service SFTP setup](/integrations/reference/sftp/self-service-setup);
  pull behaviour (files are pulled once, default schedule every 30 minutes) in
  [SFTP: Behaviour of a customer-hosted connection](/integrations/reference/sftp/sftp#behaviour-of-a-customer-hosted-connection).

Each file has a fixed location and naming pattern on S3, given in the
specification pages. The base path is:

```
s3://techwolf-<customer>/production/external/input/integrations/file_based/
```

<Note>
  During the testing phase, the `staging` environment path is used instead:
  `s3://techwolf-<customer>/staging/external/input/integrations/file_based/`.
</Note>

Files must be written into the per-file subfolders given in the specification
pages (for example `file_based/TechWolf_Job_Information/`), not into the root of
the file-based path. For the wider S3 layout and the recommended folder
structure on SFTP servers, see
[File Organization](/integrations/reference/file-structure).

## Reading the field tables

Each field is annotated in the **Key** column to indicate its role:

* **Primary key**: the external ID of the file; uniquely identifies the entity.
* **Foreign key**: references the primary key of another entity, linking the
  two.

Fields without a Key annotation are TechWolf data fields used for skill
inference or context, not identifiers.

The **Required** column is `Required` or `Recommended`. A `Required` field must
have a value on every row. A `Recommended` field can be left empty, but send it
wherever the source system holds a value.

The **Max chars** column gives the longest value the SkillEngine API accepts for
that field. A row whose value exceeds the limit is rejected on ingest, so
truncate or drop over-long values in the export rather than letting the
integration reject the record. `No limit` means the API sets no maximum; `n/a`
means the field is not free text (a date, a boolean or a number) and the Notes
column describes the format instead.

Limits are applied to the value that reaches the API, so they are measured after
any transformation configured on the integration (for example the
[Slugify](/integrations/datasource-integrations/data-integrator/how-to-set-up-integration)
step applied to identifiers).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.