A data pipeline reads your data where it already lives — a database, a data lake, a file — shapes it into Simba’s modelling format, and saves each result as a versioned dataset you can build models on. You build a pipeline once in the app; after that it can be run by you, by an agent over the API or Simba MCP, or on a schedule (see Refresh data on a schedule).
| Source | What it reads | You provide |
|---|---|---|
| CSV upload | A file you upload | The file |
| PostgreSQL | A table or a read-only SQL query | Host, port, database, user, password, and a table or query |
| Amazon Athena | A Glue Data Catalog table or an Athena query — a good fit for a Fivetran Managed Data Lake | AWS region, an S3 location for query results, a database and table (or a query), and an AWS access key |
| S3 Iceberg | An Apache Iceberg table on S3 through the AWS Glue catalog, read directly with no query engine | AWS region, the Glue namespace and table, and an AWS access key |
The builder lists only the sources the service can run. Other connectors (Snowflake, Databricks, weather and economic indicators) appear on deployments that include their drivers; where a driver is missing, the source is hidden rather than failing when it runs. A pipeline that already uses a hidden source still opens, and a run of it stops with a message naming the missing driver.
***; send *** back to keep the saved value. Credentials are stored encrypted and are never copied into a pipeline’s saved versions.In the app, open Data Pipelines and create a pipeline. Drag a source onto the canvas, then add steps — filter, merge, aggregate, calculate, clean — and mark the step whose output is the dataset. Each step previews its output as you edit, so you can see the columns before anything is saved.
Press Run pipeline. The run happens on the server: the page follows it, shows how long it has been running, and you can leave the page — the run carries on. When it finishes, each step shows its columns, the new version appears in the Runs menu, and its output is available in your datasets for model building. If it fails, the step that failed is marked and the message says why and what to change.
Only one run of a pipeline happens at a time. Pressing Run while another run is going — started in another tab, by an agent, or by the schedule — follows that run instead of starting a second one.
The Runs menu lists every run: when, who started it (you, the schedule, or an API key by name), the version it saved, and why a run failed.
Both need an API key with the create:models scope, and see only the key owner’s pipelines. pipeline_ref is a pipeline’s hash or id (from list_pipelines).
POST /api/v1/pipelines/{pipeline_ref}/runs → 202 {"run_id": 41, "status": "queued"}
GET /api/v1/pipelines/{pipeline_ref}/runs/{run_id} → {"status": "succeeded", "version_id": 212, ...}
Over MCP:
run_pipeline(pipeline_ref="ab12cd34ef")
get_pipeline_run(pipeline_ref="ab12cd34ef", run_id=41)
run_pipeline answers at once; poll get_pipeline_run every few seconds until status is succeeded or failed. An optional start_date / end_date (YYYY-MM-DD) limits the source steps to that date range. A successful run returns the version_id it saved — pass it as pipeline_version_id to get_recipe_draft_template to build a model on the refreshed data.
| Status | Meaning |
|---|---|
queued |
Waiting for a worker, usually a second or two |
running |
Reading and shaping the data |
succeeded |
A new version was saved (version_id) |
failed |
See error_code and error |
error_code |
What to do |
|---|---|
execution_failed |
A source or step failed; error names it. Fix it in the builder and run again |
no_output |
The date range returned no rows; widen it |
timeout |
The run passed its 30-minute limit; narrow the query or the date range |
interrupted, not_started |
The run was lost on the server; start it again |
not_queued |
The run could not be queued; try again shortly |
owner_blocked |
The pipeline owner’s account is blocked, so its runs do not start |
Starting a run while one is going returns 409 run_in_progress with that run’s run_id: poll it rather than starting another.
To go from a saved version to a fitted model, list the pipeline’s versions with list_pipeline_versions, take a draft template for the version you want, and either fit it directly with create_model or publish it as a study recipe that records the pipeline, version and hash it came from. The steps, the HTTP routes and what a version pins are in Build a model from a pipeline version.
The exact MCP parameters are in the simba-mcp tool reference.
A pipeline whose output is one row per campaign per day can also feed Simba’s campaign facts: see Campaign data.