> For the complete documentation index, see [llms.txt](https://docs.soda.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.soda.io/soda-ai/soda-mcp/mcp-tools.md).

# MCP tools

The soda-mcp server exposes Soda Cloud's user-callable Public API v4 surface as Model Context Protocol tools, **letting your agent** **manage your data quality**, from datasets and contracts to secrets and runners, directly. The catalog below is grouped by resource type and generated from the server's tool definitions, so it stays in sync with what the server actually offers.

Each tool maps to a documented operation within Soda v4, with MCP-specific transformations and exclusions noted where they apply.

<details>

<summary>Attributes - labels for organizing datasets, checks, and columns</summary>

Attributes categorize and organize resources, such as marking checks as critical for notification rules.

* `list_attributes` — List attribute definitions available in your Soda Cloud organization for datasets, checks, columns, and data standards.
* `create_attribute` — Create a new attribute definition for datasets, checks, columns, or data standards.
* `update_attribute` — Update an attribute's label, description, or allowed values.
* `delete_attribute` — Delete an attribute definition. The attribute is removed from any resources where it is set.

</details>

<details>

<summary>Checks - data quality validations from multiple sources</summary>

Checks are Soda Cloud data quality validations and can originate from Soda products, dbt, data contracts, data standards, or metric monitoring. Contract-authored checks can use fixed thresholds; metric-monitoring checks use dynamic thresholds. There is no generic public create-check endpoint.

* `list_checks` — List checks in your Soda Cloud organization, including current status and latest result, associated datasets, agreements, and linked incidents.
* `list_check_results` — List historical results for one check, including result IDs, evaluation statuses, and data timestamps.
* `delete_check` — Delete a specific check by its ID. Checks are the general Soda Cloud validation resource and can originate from contracts, data standards, Soda products, dbt, or metric monitoring.

</details>

<details>

<summary>Datasets - tables tracked in Soda Cloud</summary>

* `list_datasets` — List datasets in your Soda Cloud organization, including data source, incidents, attributes, health status, and contract metadata.
* `get_dataset` — Get a single dataset by dataset ID.
* `get_dataset_by_dataset_qualified_name` — Get a single dataset by dataset qualified name (DQN), e.g. \[datasource\_name]/\[database\_name]/\[schema\_name]/\[dataset\_name].
* `update_dataset` — Update dataset properties (label, tags, attributes, owners, profiling, metric monitoring, diagnostics warehouse, compute warehouse, time partition, data delay).
* `list_dataset_columns` — List active columns of a dataset with current assigned column attribute values.
* `set_column_attributes` — Set attribute values on columns of a dataset.
* `delete_dataset` — Delete a dataset. This action is permanent and cannot be undone.
* `list_dataset_roles` — List available dataset roles. Use dataset roles to manage access to individual datasets.
* `create_dataset_role` — Create a custom dataset role.
* `update_dataset_role` — Update the name or permissions of a custom dataset role.
* `delete_dataset_role` — Delete a custom dataset role. This action is permanent and cannot be undone.
* `get_dataset_compute_warehouse` — Get the compute warehouse configuration for a dataset.
* `get_dataset_diagnostics_warehouse` — Get diagnostics warehouse information for a dataset.
* `get_dataset_profiling` — Get profiling information for a dataset (structure, statistical summaries, data characteristics).
* `list_dataset_responsibilities` — List user and user group permissions assigned to a dataset, and their associated roles.
* `update_dataset_responsibilities` — Update user/group permissions and their associated roles for a dataset.

</details>

<details>

<summary>Monitors - ML-based dynamic-threshold metrics</summary>

Monitors track metrics over time with anomaly detection. Dataset monitors cover built-in metrics, while column and custom SQL monitors are user-defined.

* `get_dataset_metric_monitoring` — Get metric monitoring configuration for a dataset, including all three monitor types and the dataset-level scan schedule.
* `create_column_metric_monitor` — Create a user-defined column metric monitor for a dataset.
* `update_column_metric_monitor` — Update an existing user-defined column metric monitor (e.g. sensitivity, thresholds, grouping, exclusion zones).
* `delete_column_metric_monitor` — Delete a user-defined column metric monitor. This action is permanent and cannot be undone.
* `create_custom_sql_monitor` — Create a user-defined custom SQL monitor for a dataset.
* `update_custom_sql_monitor` — Update an existing user-defined custom SQL monitor (e.g. SQL query, sensitivity, thresholds, grouping, exclusion zones).
* `delete_custom_sql_monitor` — Delete a user-defined custom SQL monitor. This action is permanent and cannot be undone.
* `run_historical_metric_collection` — Trigger a historical metric collection scan for a dataset. The v4 API calls this a historical metric collection scan; in user-facing terms it backfills metric history so monitors (ML-based anomaly detection) have a baseline to learn from.

</details>

<details>

<summary>Contracts - per-dataset sources of truth for contract-authored checks</summary>

A data contract defines its contract-authored checks for one dataset. See Editing contracts for the safe edit workflow.

* `list_contracts` — List data contracts in your Soda Cloud organization. A data contract is the declarative source-of-truth for contract-authored checks on a dataset.
* `create_contract` — Create a new data contract on a dataset. The contract is initialized with the YAML contents you supply; the full contract YAML is normally installed later via `edit_contract`.
* `get_contract` — Retrieve a specific data contract by ID, including its YAML content, with an optional contract-editing mode.
* `edit_contract` — Edit an existing data contract's YAML using its contract ID and the actual `edit_revision` from a completed `get_contract(for_editing=true)` result. This is not a general editor for check YAML: data-standard YAML, including threshold patches and previews, uses `update_data_standard`. Never substitute a data-standard ID or invent a revision. Operations-mode diffs compare the normalized baseline with the patched candidate; contents-mode diffs compare the exact fetched and caller-supplied strings.
* `get_contract_editing_knowledge` — Get packaged contract-editing knowledge, returning the authoring guide by default.
* `list_contract_versions` — List published versions of a specific data contract. Each publish creates a new version; this tool returns the version history.
* `verify_contract` — Trigger a contract verification scan for the given contract. Runs the checks defined in the contract against the dataset.
* `create_skeleton_contract` — Trigger async generation of a skeleton contract for a dataset, derived from its warehouse schema. Use this to bootstrap a contract's YAML body when no contract has been published yet.
* `get_skeleton_contract_status` — Get the status of a skeleton contract generation operation started via `create_skeleton_contract`.
* `generate_contracts` — Trigger async AI-powered generation of full contracts for one or more datasets.
* `get_contract_generation_status` — Get the status of a contract generation operation started via `generate_contracts`.

</details>

<details>

<summary>Data Standards - organization-wide YAML policies applied by scope</summary>

A data standard applies checks to every dataset matching its scope, unlike a contract, which belongs to one dataset. Targeted YAML editing preserves supported public configuration and validates the whole candidate locally. See Editing data standards for modes, preview, and publication boundaries.

* `list_data_standards` — List data standards in your Soda Cloud organization. A data standard is an organization-wide YAML policy whose checks are automatically applied to every dataset matching its scope.
* `get_data_standards_activity` — Get an organization-wide rollup of data standards activity, including counts of active and total standards, matched datasets, and aggregated check results.
* `get_data_standard` — Retrieve a specific data standard by ID, including its YAML contents, attributes, and scope.
* `get_data_standard_authoring_knowledge` — Get packaged knowledge for creating and updating data standards and previewing their scope; the guide is the default.
* `list_data_standard_checks` — List the aggregated checks generated by a specific data standard across the datasets it matches.
* `list_data_standard_datasets` — List the datasets currently matched by a specific data standard's scope.
* `evaluate_data_standard_scope` — Re-evaluate which datasets currently match a data standard's scope and return their dataset IDs.
* `preview_data_standard_scope` — Preview which datasets a candidate data standard scope matches while leaving saved standards unchanged.
* `create_data_standard` — Create an organization-wide data standard whose YAML checks apply to every dataset matching its scope.
* `update_data_standard` — Edit a data standard's YAML checks or configuration by `data_standard_id`.
* `update_data_standard_status` — Transition a data standard to a new status without modifying its contents, attributes, scope, owners, or schedule. Use this to pause or activate an existing standard.
* `delete_data_standard` — Delete a data standard.
* `execute_data_standards` — Trigger a scan that runs the active data standards linked to a dataset. Identify the dataset by `datasetId`.
* `dry_run_data_standard` — Test an unsaved data standard (a dry run) by running its checks once against one dataset's live data, without creating or modifying any data standard. Provide the standard YAML in `contents` and the target dataset in `datasetId`.

</details>

<details>

<summary>Datasources - warehouse connections that own datasets</summary>

A datasource is a configured warehouse connection, such as Snowflake or Postgres, that owns datasets.

Reading datasource configuration YAML can return stored literal credentials to the MCP host; those values enter model context when the host or embedder forwards the tool result. Soda secret references remain `${secret.NAME}` placeholders. Prefer Soda secrets and the Soda Cloud Web UI for future credential changes. This server does not redact configuration-read results or control host or model-provider retention.

* `list_datasources` — List datasources in your Soda Cloud organization. A datasource is a configured connection to a warehouse (e.g. Snowflake, Postgres) that owns datasets.
* `run_datasource_discovery` — Trigger a discovery scan on a datasource. Detects new tables and schemas in the warehouse on-demand (instead of waiting for the scheduled cron).
* `create_datasource` — Create a new datasource from a YAML configuration document. The datasource type (e.g. Snowflake, Postgres) is extracted from the YAML contents.
* `get_datasource` — Get datasource metadata by ID: ID, name, label, type, owner, and creation/update timestamps.
* `get_datasource_configuration_file_contents` — Get the current datasource connection configuration YAML by ID. Use this read before a targeted edit: supplying `configurationFileContents` to `update_datasource` replaces the whole YAML, so preserve every unchanged returned value.
* `update_datasource` — Update a datasource. Top-level fields are PATCH-like: omitted fields keep their existing values. Supplying `configurationFileContents` instead replaces the entire YAML; for a targeted YAML edit, first call `get_datasource_configuration_file_contents` and preserve every unchanged returned value, including existing literals.
* `delete_datasource` — Delete a datasource and all its associated resources (datasets, checks, scans, incidents).
* `list_datasource_roles` — List available datasource roles. Use datasource roles to manage access to individual datasources.
* `create_datasource_role` — Create a custom datasource role.
* `update_datasource_role` — Update the name or permissions of a custom datasource role.
* `delete_datasource_role` — Delete a custom datasource role. This action is permanent and cannot be undone.
* `get_datasource_diagnostics_warehouse` — Get the diagnostics warehouse configuration for a datasource. The diagnostics warehouse collects scan-related data (failed rows, scan results) and forwards it to the customer's warehouse for storage and analysis.
* `update_datasource_diagnostics_warehouse` — Update the diagnostics warehouse configuration for a datasource. The diagnostics warehouse collects scan-related data (failed rows, scan results) and stores it in your warehouse for analysis.
* `list_datasource_responsibilities` — List user and user group permissions assigned to a datasource, and their associated roles.
* `update_datasource_responsibilities` — Update user/group permissions and their associated roles for a datasource.
* `test_datasource_connection` — Trigger an async connection test for a datasource configuration. Use this to validate a YAML + runner combination without creating a datasource, e.g. before calling `create_datasource`.
* `get_datasource_connection_test_status` — Get the status of an async datasource-connection-test operation started via `test_datasource_connection`.
* `onboard_discovered_datasets` — Trigger async onboarding of one or more discovered datasets for a datasource. Onboarding promotes a discovered-but-not-yet-tracked dataset into a regular Soda dataset that can carry contracts, checks, and monitors.
* `get_onboard_discovered_datasets_status` — Get the status of an async dataset-onboarding operation started via `onboard_discovered_datasets`.

</details>

<details>

<summary>Discovered Datasets - tables found by discovery scans</summary>

Discovered datasets have been found by Soda but have not yet been onboarded as regular Soda datasets.

* `list_discovered_datasets` — List datasets Soda has discovered in your data sources. Discovered datasets have been found during a discovery scan but may not yet be onboarded as first-class Soda datasets.

</details>

<details>

<summary>Runners - execute scans and datasource operations</summary>

Self-hosted runners require one-time API key credentials created through `create_runner`.

* `list_runners` — List Soda runners in your organization.
* `create_runner` — Create API key credentials for a new self-hosted Soda runner deployment.
* `get_runner` — Get a runner by ID, including online status, runner type, version information, and last-seen timestamp.
* `delete_runner` — Delete a self-hosted runner by ID.

</details>

<details>

<summary>Secrets - encrypted credentials referenced from datasource YAML</summary>

Datasource YAML can reference stored credentials as `${secret.NAME}`. The Soda Cloud Web UI is the safer place for a person to enter a secret because it keeps plaintext out of chat history and the MCP client/server flow.

`create_secret` and `update_secret` accept plaintext locally and follow the same authorization rule as other writes: a well-specified request needs no separate warning-and-confirmation exchange. The Web UI advice is informational. The server encrypts the value before transmitting it to Soda Cloud.

* `list_secrets` — List secrets in your organization.
* `create_secret` — Create a new secret that can be referenced in datasource YAML as `${secret.NAME}`.
* `delete_secret` — Delete an existing secret.
* `update_secret` — Update the value of an existing secret. The secret name cannot be changed.

**Secrets threat model**

* Plaintext exists in local agent process memory and on the local stdio pipe, but is never logged or sent to Soda Cloud unencrypted.
* The server uses AES-256-GCM and wraps the key with Soda Cloud's RSA public key. It caches that public key for about an hour and refreshes it once after an `encryption_failed` response.
* This is defense in depth over TLS. It does not protect against a compromised MCP client or local agent process.

</details>

<details>

<summary>Incidents - triage state on data quality failures</summary>

* `update_incident` — Update an incident's title, severity, status, description, resolution notes, lead, or linked check results.

</details>

<details>

<summary>Scans - observe and control scan executions</summary>

Contract, datasource, monitor, and data-standard tools trigger scans. These tools inspect or cancel those scan executions.

* `get_scan_status` — Get the current state of a scan. Use this to monitor scan progress during execution.
* `get_scan_logs` — Get log details for a completed scan. Use this to investigate issues with scan execution.
* `cancel_scan` — Cancel a running scan.

</details>

<details>

<summary>Users - organization users and groups</summary>

* `list_users` — List users in your Soda Cloud organization. Search by first name, last name, or email.
* `create_user` — Invite one or more users to your Soda Cloud organization.
* `disable_user` — Disable a user in your Soda Cloud organization.
* `list_user_groups` — List user groups in your Soda Cloud organization, including lists of members. Fuzzy search on group name.
* `get_user_group` — Get a single user group by ID, including its members. Returns the full group details (non-paginated).
* `create_user_group` — Create a new user group. Use this tool whenever the user asks to create a user group (typically naming a new group), even when the prompt also lists initial members.
* `update_user_group` — Update an existing user group's members. Use this tool only when the user references an existing group (by ID or by name); for a brand-new group, use `create_user_group` instead.
* `delete_user_group` — Delete a user group by ID. This action is permanent and cannot be undone.

</details>

<details>

<summary>Licensing - SPU consumption</summary>

SPUs (Soda Processing Units) are the consumption unit of SPU-based licensing. Consumption accumulates per day. An organization on any other licensing model is refused with a `400 organisation_not_on_spu_plan` rather than reading zero, so a zero total always means nothing was consumed over the period asked about.

* `get_spu_consumption` — Get the number of SPUs (Soda Processing Units) your organization consumed over a period.

</details>

***

{% hint style="info" %}
You are **not logged in to Soda** and are viewing the default public documentation. Learn more about [Licensing & documentation access](/reference/documentation-access-and-licensing.md).

If you do have a Soda license, make sure to **log in to Soda Cloud in this same browser**.
{% endhint %}


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.soda.io/soda-ai/soda-mcp/mcp-tools.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
