Skip to main content
This page covers the parts shared by every file-based datasource specification in this section: how files travel from the customer to TechWolf, where they land on S3, and how to read the field tables on the Jobs, Employees and Courses pages. For the broader picture of how file-based integrations work and how they compare to API-based connectors, see Key concepts.

Technical data flow

The files described in the specification pages end up in the customer’s TechWolf AWS S3 bucket, under the file-based ingest prefix. The file-based integration reads them from there and syncs the data into the TechWolf SkillEngine API. How the files reach S3 depends on the delivery method agreed during implementation:
  • TechWolf-hosted SFTP server. The customer (or their source system) pushes the files to the TechWolf SFTP server. Each SFTP user’s root directory maps to a directory in TechWolf’s S3 infrastructure, so an upload lands on S3 directly. TechWolf provides the credentials. Server addresses and static IP addresses for allowlisting are listed under SFTP: TechWolf-hosted SFTP server.
  • Customer-hosted SFTP server. The customer exports the files to their own SFTP server. TechWolf connects to it on a schedule, pulls the files, and writes them to S3 according to the folder mappings configured in the Console. Setup is described in Self-service SFTP setup; pull behaviour (files are pulled once, default schedule every 30 minutes) in SFTP: Behaviour of a customer-hosted connection.
Each file has a fixed location and naming pattern on S3, given in the specification pages. The base path is:
During the testing phase, the staging environment path is used instead: s3://techwolf-<customer>/staging/external/input/integrations/file_based/.
Files must be written into the per-file subfolders given in the specification pages (for example file_based/TechWolf_Job_Information/), not into the root of the file-based path. For the wider S3 layout and the recommended folder structure on SFTP servers, see File Organization.

Reading the field tables

Each field is annotated in the Key column to indicate its role:
  • Primary key: the external ID of the file; uniquely identifies the entity.
  • Foreign key: references the primary key of another entity, linking the two.
Fields without a Key annotation are TechWolf data fields used for skill inference or context, not identifiers. The Required column is Required or Recommended. A Required field must have a value on every row. A Recommended field can be left empty, but send it wherever the source system holds a value. The Max chars column gives the longest value the SkillEngine API accepts for that field. A row whose value exceeds the limit is rejected on ingest, so truncate or drop over-long values in the export rather than letting the integration reject the record. No limit means the API sets no maximum; n/a means the field is not free text (a date, a boolean or a number) and the Notes column describes the format instead. Limits are applied to the value that reaches the API, so they are measured after any transformation configured on the integration (for example the Slugify step applied to identifiers).