# Guides

<div data-with-frame="true"><figure><img src="/files/JOvm6LbwFvugzm514bZ4" alt=""><figcaption></figcaption></figure></div>

**Paradime is the AI-native platform for analytics engineers — built to replace dbt Cloud™ in the AI era.**

Think Cursor for Data. Write, run, document, and fix your dbt pipelines with AI agents that understand your project — not just your prompt.

Ranked **#1 on ADE-Bench** with an 88.37% score — ahead of dbt Labs, Claude Code, and Cortex CLI.

### What's inside

Paradime brings four core products into one workspace:

<table data-view="cards"><thead><tr><th align="center"></th><th align="center"></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td align="center"><strong>🖥️ Code IDE</strong></td><td align="center"><strong>Your dbt Development Environment</strong></td><td>A purpose-built IDE for analytics engineers. Write dbt models, run queries, explore your schema, and manage your git workflow — all without leaving Paradime.</td><td><a href="/pages/9gQ6Hok87Ezvjljn06WV">/pages/9gQ6Hok87Ezvjljn06WV</a></td></tr><tr><td align="center"><strong>🦖 DinoAI</strong></td><td align="center"><strong>Your AI Data Engineering Agent</strong></td><td>DinoAI isn't a chatbot. It's a warehouse-aware agent that reads your dbt project, understands your schema, and autonomously codes, refactors, documents, and deploys — straight from your IDE.</td><td></td></tr><tr><td align="center"><strong>⚡ Bolt</strong></td><td align="center"><strong>Production-grade Data Orchestration</strong></td><td>Smart scheduling, CI/CD pipelines, instant error alerts, and one-click import from dbt Cloud™. Configure once and let your pipelines run on autopilot.</td><td></td></tr><tr><td align="center"><strong>📡 Radar</strong></td><td align="center"><strong>Warehouse FinOps</strong></td><td>Reduce your Snowflake and BigQuery spend without touching a single model. Radar gives CFOs, COOs, and data leaders full visibility into credit consumption — and frees up budget for AI use cases.</td><td></td></tr></tbody></table>

### Key results from teams using Paradime

| Metric                    | Result     |
| ------------------------- | ---------- |
| Development speed         | 73% faster |
| MTTR on pipeline failures | 60% lower  |
| Warehouse cost reduction  | 20%+       |
| Uptime SLA                | 99.80%     |

{% hint style="success" icon="comet" %}
*Trusted by Tide, Motive, Customer.io, Sephora, and hundreds of other data teams.*
{% endhint %}


# Paradime 101

Paradime 101 is your comprehensive onboarding guide to Paradime's powerful analytics engineering platform. Designed specifically for new Paradime users, this guide will quickly familiarize you with Paradime's key features and best practices, enabling you to add value to your organization's data processes right from the start.

**Estimated completion time:** 3-4 hours

***

### What you'll learn

<table data-view="cards"><thead><tr><th></th><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td>🏗️ <strong>Getting Started with your Paradime workspace</strong></td><td></td><td>Estimated time: 45 minutes<br><br>Learn how to set up and manage your Paradime workspace, including creating a workspace, connecting data warehouses, and managing users.</td><td><a href="/pages/BkOIZcQ7Fway6AcUnJYR">/pages/BkOIZcQ7Fway6AcUnJYR</a></td></tr><tr><td>💻 <strong>Getting Started with the Paradime IDE</strong></td><td></td><td>Estimated time: 90 minutes<br><br>Discover how to use the Paradime IDE for dbt™ development, including project setup, model creation, data exploration, and leveraging DinoAI for accelerated workflows.</td><td><a href="/pages/qsrMdBdV5pHhUOwfRlSD">/pages/qsrMdBdV5pHhUOwfRlSD</a></td></tr><tr><td>⚡ <strong>Managing dbt™ Schedules with Bolt</strong></td><td></td><td>Estimated time: 45 minutes<br><br>Master the use of Bolt for efficiently scheduling and managing your dbt™ jobs in production environments.</td><td><a href="/pages/YQDWv1QRdxtLjMkqnP7R">/pages/YQDWv1QRdxtLjMkqnP7R</a></td></tr></tbody></table>

***

### How to use this guide

* Follow the sections in order for a comprehensive learning experience.
* Use the "Related Documentation" on each page to jump to specific topics if you need targeted information.
* Implement as you learn to reinforce your learning.
* Refer back to relevant sections as you work on your own projects in Paradime.

***

Let's begin your journey to becoming a Paradime expert and enhancing your organization's data capabilities! Start with [Getting Started with your Paradime workspace](/app-help/guides/paradime-101/getting-started-with-your-paradime-workspace).


# Getting Started with your Paradime Workspace

This first section of the Paradime 101 Guide will walk you through setting up and managing your workspace, the foundation for all your dbt™ development activities. A well-configured workspace is crucial for efficient collaboration and streamlined analytics engineering processes.

**Estimated completion time:** 40 minutes

***

### What you'll learn

By the end of this guide, you'll have a fully configured workspace ready for collaborative analytics engineering. You'll learn about:

<table data-view="cards"><thead><tr><th></th><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td>🏗️ <strong>Creating a Paradime Workspace</strong></td><td></td><td>Learn how to create your Paradime workspace, connect a Git repository, and set up initial data warehouse connections.<br><br><em>Estimated time: 10 minutes</em></td><td><a href="/pages/EF657oVXCXaTNCJUqXTD">/pages/EF657oVXCXaTNCJUqXTD</a></td></tr><tr><td>🔌 <strong>Setting Up Data Warehouse Connections</strong></td><td></td><td>Dive deeper into configuring multiple data warehouse connections for development, testing, and production environments.<br><br><em>Estimated time: 10 minutes</em></td><td><a href="/pages/IZu4vDaG94nR9oplKDwI">/pages/IZu4vDaG94nR9oplKDwI</a></td></tr><tr><td>⚙️ <strong>Managing Workspace Configurations</strong></td><td></td><td>Explore essential workspace configurations, including dbt™ version management, environment variables, and API keys.<br><br><em>Estimated time: 10 minutes</em></td><td><a href="/pages/iBy6R1ITBJniQ78Hbljw">/pages/iBy6R1ITBJniQ78Hbljw</a></td></tr><tr><td>👥 <strong>Managing Users in the Workspace</strong></td><td></td><td>Learn how to effectively manage user access, roles, and permissions in your Paradime workspace.<br><br><em>Estimated time: 10 minutes</em></td><td><a href="/pages/cG0eBVac6ivmN9HKjdRe">/pages/cG0eBVac6ivmN9HKjdRe</a></td></tr></tbody></table>

Let's dive in!


# Creating a Workspace

In this guide, you'll learn all the essential steps to create a Paradime workspace, the foundation for your dbt™ development journey. A well-configured workspace is crucial for efficient collaboration and streamlined analytics engineering processes.

**Estimated completion time:** 10 minutes

{% hint style="warning" %}

#### Prerequisites

* An [**admin role**](/app-help/documentation/settings/users/role-based-access-control) In Paradime to create workspaces
* Repository admin access to add a deploy key for your [**Git repository**](/app-help/documentation/settings/git-repositories/importing-a-repository)
* Access credentials for your [**data warehouse**](/app-help/documentation/settings/connections/development-environment)
  {% endhint %}

### What You'll Learn

In this guide, you'll learn how to:

1. [Create a Paradime workspace](#id-1.-creating-a-paradime-workspace)
2. [Connect a Git Repository](#id-2.-connect-a-git-repository-to-the-workspace)
3. [Add a Data warehouse connection](#id-3.-add-a-data-warehouse-connection)

***

### Video Tutorial

The following video provides step-by-step instructions for creating a workspace, connecting a Git repository, and setting up a data warehouse connection:

{% embed url="<https://www.youtube.com/watch?feature=youtu.be&v=ONAqHL8t6wQ>" %}

### 1. Creating a Paradime Workspace

To create a workspace in Paradime:

1. Navigate to your Platform settings.
2. Click the `New Workspace` button to begin setting up your workspace.
3. Provide a name for your workspace.
4. Choose whether you want other people in your organization to access it without an invite (with a business user role). This option makes the workspace visible in the workspaces list to users, even if they're not yet part of it.

***

### 2. Connect a Git Repository to the Workspace

Next, connect a Git repository to Paradime:

* This can be an empty repository or an existing dbt™ project.
* Upon adding a repository SSH URI, Paradime will generate a deploy key.
* Use this deploy key to grant Paradime write access to the repo, allowing users to create commits and push branches from the Paradime IDE.

{% hint style="info" %}
**Note:** You must be a repository admin to add a deploy key. Paradime supports [GitHub](/app-help/documentation/settings/git-repositories/importing-a-repository/github), [GitLab](/app-help/documentation/settings/git-repositories/importing-a-repository/gitlab), [Azure Repos](/app-help/documentation/settings/git-repositories/importing-a-repository/azure-devops), and [Bitbucket](/app-help/documentation/settings/git-repositories/importing-a-repository/bitbucket).
{% endhint %}

***

### 3. Add a Data Warehouse Connection

Finally, add a data warehouse connection during the workspace onboarding:

* This enables developing and running dbt models via the Paradime IDE.
* Paradime supports many [warehouse providers](/app-help/documentation/settings/connections/development-environment), including:
  * [Snowflake](/app-help/documentation/settings/connections/development-environment/snowflake)
  * [BigQuery](/app-help/documentation/settings/connections/development-environment/bigquery)
  * [Redshift](/app-help/documentation/settings/connections/development-environment/redshift)
  * [Databricks](/app-help/documentation/settings/connections/development-environment/databricks)
* Multiple authentication methods are available based on your needs.

***

{% hint style="info" %}
Related Documentation:

* [Setting up Data Warehouse Connections](/app-help/documentation/settings/connections/development-environment)
* [Managing workspace configurations](/app-help/documentation/settings/workspaces)
* [Importing a Git Repository](/app-help/documentation/settings/git-repositories/importing-a-repository)
  {% endhint %}

***

### Summary

You've successfully created your Paradime workspace, which serves as the foundation for your dbt™ development. Here's what you've accomplished:

1. Created a new workspace in Paradime
2. Connected a Git repository to your workspace
3. Set up an initial data warehouse connection

In the next guide, we'll explore [how to configure additional data warehouse connections](/app-help/guides/paradime-101/getting-started-with-your-paradime-workspace/setting-up-data-warehouse-connections) for development, testing, and production environments.


# Setting Up Data Warehouse Connections

In this guide, you'll learn how to set up and manage multiple data warehouse connections within your Paradime workspace. This process is crucial for configuring connections for both development in the [Paradime Code IDE](/app-help/documentation/code-ide) and running dbt™ in production with the [Bolt Scheduler](/app-help/documentation/bolt).

**Estimated completion time:** 10 minutes

{% hint style="warning" %}

#### Prerequisites

* An [**admin role**](/app-help/documentation/settings/users/role-based-access-control) In Paradime to create workspaces
  {% endhint %}

### What you'll learn

In this guide, you'll learn how to:

1. [Add a Bolt Scheduler connection](#id-2.-add-a-bolt-scheduler-connection)
2. [Add a TurboCI connection](#id-3.-add-a-turboci-connection)
3. [Manage your connections](#id-4.-manage-your-connections)
4. [Use targets with dbt™ commands](#id-5.-using-targets-with-dbt-tm-commands)

***

### Video Tutorial

The following video provides step-by-step instructions for adding a scheduler environment (bolt scheduler) and a Turbo CI connection.

{% embed url="<https://youtu.be/DdPo-dMbkBE>" %}

***

### 1. Add a Bolt Scheduler Connection

This connection will be used as the default target for your Bolt schedules to run dbt™ in production.

1. Navigate to your account settings
2. Locate the `Connections` tab.
3. Click `Add New` button in the Scheduler section.
4. Select your data warehouse provider (e.g., Snowflake)
5. Configure the connection details.
6. Click 'Test Connection' to verify the setup.

💡 For visual guidance, follow [video tutorial](https://youtu.be/DdPo-dMbkBE?si=dyCMB9KyP8K3BFYG\&t=32).

***

### 2. Add a TurboCI Connection

This connection will be used specifically for running dbt™ with [TurboCI](/app-help/documentation/bolt/ci-cd/turbo-ci) to build and test your dbt™ models when opening a Pull Request.

1. Follow steps 1-3 from the above section, [Add a Bolt Scheduler Connection](#id-2.-add-a-bolt-scheduler-connection)
2. Configure the connection details:
   * Provide a dbt™ profile name (This should match with the profile name set in your `dbt_project.yml`).
   * Set the target as `CI`.
   * Enter your data warehouse credentials (similar to the Scheduler prod connection).
   * Change the schema name to `dbt_ci` (or your preferred CI schema).
   * Configure thread settings as needed.
3. Click `Test Connection` to verify the setup.

💡 For visual guidance, follow [video tutorial](https://youtu.be/DdPo-dMbkBE?si=D8M0ljlb8Py99ORh\&t=102).

***

### 3. Manage Your Connections

After setting up your connections, you'll have:

* A `prod` target connection for orchestrating production jobs.
* A `ci` target connection for running dbt™ with [TurboCI](/app-help/documentation/bolt/ci-cd/turbo-ci).

These connections enable you to separate your development, testing, and scheduler environments effectively within Paradime.

***

### 4. Using Targets with dbt™ Commands

After setting up multiple connections, it's important to understand how to direct your dbt™ commands to use a specific target. This allows you to switch between multiple environments easily.

#### Default Behavior

By default, when you run a dbt™ command without specifying a target:

* In the Code IDE: It will use the default connection set for the Code IDE.
* In the Scheduler: It will use the default connection set for the Scheduler.

Example:

```bash
dbt run
```

It will use the default target for your current environment.

#### Specifying a Target

To use a specific target, append the `--target` argument to your dbt command:

```bash
dbt <command> --target <target_name>
```

For example, to run dbt using the 'CI' target we set up earlier:

```bash
dbt run --target ci
```

This command will execute the dbt run using the connection details specified in the `ci` target, regardless of the default setting.

***

{% hint style="info" %}

#### Related Documentation

* [Production Environments in Paradime](/app-help/documentation/settings/connections/scheduler-environment)
* [CI/CD In Paradime](/app-help/documentation/bolt/ci-cd)
  {% endhint %}

***

### Summary

You've now set up multiple data warehouse connections, learned how to manage them, and understand how to use different targets with dbt™ commands. These skills will allow you to efficiently work across development, testing, and production environments in Paradime.

Next, we'll dive into [Managing Workspace Configurations](/app-help/guides/paradime-101/getting-started-with-your-paradime-workspace/managing-workspace-configurations) to further customize your Paradime environment.


# Managing workspace configurations

In this guide, you'll learn how to manage essential configurations for your Paradime workspace. Proper configuration management ensures consistency, security, and flexibility across your analytics engineering projects.

**Estimated completion time:** 10 minutes

{% hint style="warning" %}

#### Prerequisites

* An [**admin role**](/app-help/documentation/settings/users/role-based-access-control) In Paradime to create workspaces
  {% endhint %}

### What you'll learn

* [Manage dbt™ versions for your workspace](#id-1.-managing-dbt-tm-versions-in-paradime-workspaces)
* [Create and Manage environment variables](#id-2.-setting-up-environment-variables)
* [Create and manage API keys](#id-3.-managing-api-keys)

***

### 1. Managing dbt™ Versions in Paradime Workspaces

Paradime allows you to centrally manage the dbt™ version for your entire workspace, ensuring consistency across both the Code IDE and Bolt scheduler for all users. This feature is crucial for maintaining compatibility and leveraging the latest dbt™ functionalities across your projects.

{% embed url="<https://www.youtube.com/watch?v=xJFzaTTfvvs>" %}

#### View your current dbt version:

1. Navigate to your account settings.
2. Select the `Workspace` tab.
3. Under the project settings section, find the current dbt™ version configuration.

#### Upgrade or downgrade your dbt ™version

{% hint style="warning" %}
Before changing the dbt™ version, ensure that your projects are compatible with the new version to avoid any potential issues.
{% endhint %}

* Click the `Edit` button next to the current dbt™ version.
* Select the desired dbt™ version from the dropdown menu.
* Click the `Save` button to apply the changes.

Once saved, all users in the workspace will inherit this dbt™ version, which applies to both the Code IDE and the Bolt scheduler.

***

### 2. Setting Up Environment Variables

Environment variables allow you to securely store and use sensitive information or configuration details in Paradime. These variables can be managed in two different contexts:

{% hint style="info" %}

* **Bolt Schedule Environment Variables**: For production jobs (requires Admin access)
* **Code IDE Environment Variables**: For development work
  {% endhint %}

{% embed url="<https://www.youtube.com/watch?v=FPIphtub40k>" %}

#### Add a New Environment Variable

1. From any page in the Paradime application, Click **Settings**
2. Navigate to **Workspaces > Environment Variables**
3. Choose where to add your variable:
   1. **For production jobs**: In the Bolt Schedules section, click **Add New**
   2. **For development work**: In the Code IDE section, click **Add New**
4. Configure your variable:
   1. Enter Key name
   2. Enter Value
   3. Click Save icon (💾)

{% hint style="success" %}
**When Variables Take Effect**

* **Bolt Schedules**: Variables are automatically injected into your dbt™ project at runtime
* **Code IDE**: Variables are immediately available in your development environment
  {% endhint %}

#### Managing Environment Variables

* **Edit**: Click the 'Edit' button (✎) next to the variable
* **Delete**: Click the 'Delete' button (🗑️) next to the variable
* Remember to save your changes

{% hint style="warning" %}
**Important notes**

* Bolt Schedule variables require Admin access to manage
* Environment variables are environment-specific (development vs. production)
* For detailed information, see:
  * [Bolt Schedules Environment Variables](/app-help/documentation/settings/environment-variables/bolt-schedule-env-variables)
  * [Code IDE Environment Variables](/app-help/documentation/settings/environment-variables/code-ide-env-variables)
    {% endhint %}

***

### 3. Managing API Keys

API keys in Paradime provide secure, programmatic access to your workspace resources. They are essential for integrating Paradime with other tools and automating workflows. Each API key is scoped to a specific workspace, ensuring isolation and security of your data and configurations.

{% embed url="<https://youtu.be/TEeNQVhlKdA>" %}

#### Generate a New API Key

1. Navigate to your account settings.
2. Select the "Workspace" tab.
3. Scroll to API key section and click on the `+ Generate API Key`
4. Provide a name for the key (e.g., 'My New API Key').
5. Optional: Set a lifetime for the key by specifying the number of days until expiration.
6. Click `Create` to save.
7. After creation, securely store your API key, secret, and endpoint.

#### Define API Key Capabilities

Choose the capabilities you want to enable for this API key, you can find all the capabilities and access in the [API docs](/app-help/developers/generate-api-keys-legacy). Select the appropriate capabilities based on your intended use of the API key.

{% hint style="info" %}
**Important Note on API Key Scope:** API keys in Paradime are scoped to specific workspaces. This means:

* Each API key is tied to the workspace in which it was created.
* An API key generated in one workspace cannot be used to access or manage resources in another workspace.
  {% endhint %}

#### Managing API Keys

To manage existing API keys:

* View all generated keys in the API key section of your workspace settings.
* Delete keys that are no longer needed or may have been compromised.

{% hint style="info" %}
**Best Practice:** Regularly review your API keys, rotate them periodically, and ensure each key has only the necessary capabilities for its intended use.
{% endhint %}

***

{% hint style="info" %}

#### Related Documentation

* [Upgrade dbt™ core version](/app-help/documentation/settings/dbt/upgrade-dbt-core-version)
* [Environment Variables](/app-help/documentation/settings/environment-variables)
* [API Keys](/app-help/developers/generate-api-keys-legacy)
  {% endhint %}

***

### Summary

You've now learned how to manage dbt™ versions, set up environment variables, and handle API keys in your Paradime workspace. These skills will help you maintain a secure and efficient analytics engineering environment.

Next, we'll explore [Managing Users in the Workspace](/app-help/guides/paradime-101/getting-started-with-your-paradime-workspace/managing-workspace-configurations) to ensure proper access control and collaboration in your Paradime projects.


# Managing Users in the Workspace

In this guide, you'll learn how to effectively manage users in your Paradime workspace. Proper user management is crucial for maintaining security, controlling access, and facilitating collaboration in your analytics engineering projects.

**Estimated completion time:** 10 minutes

{% hint style="warning" %}

#### Prerequisites

* An [**admin role**](/app-help/documentation/settings/users/role-based-access-control) In Paradime to create workspaces
  {% endhint %}

### What you'll learn

In this guide, you'll learn how to:

1. [Access user management](#id-1.-access-user-management)
2. [Invite new users to your workspace](#id-2.-inviting-new-users)
3. [Manage user roles](#id-3.-managing-user-roles)
4. [Deactivate users when necessary](#id-4.-deactivating-users)
5. [Set up auto-join email domains](#id-5.-auto-join-email-domains)

***

### Video Tutorial

The following video provides step-by-step instructions for managing users in your Paradime workspace:

{% embed url="<https://youtu.be/_b5Jy8oA8M8>" %}

***

### 1. Access User Management

To begin managing users:

1. Navigate to your account settings
2. Select the "Team" section from the left menu.

Here, you'll find the user management panel, displaying all current active users, their roles, and their last login time.

***

### 2. Inviting New Users

To invite new users to your workspace:

1. Click on the `New User` button at the top right.
2. Choose to invite via Slack or email.
3. For email invitations:
   1. Enter the user's email address(es), separated by commas for multiple invites.
   2. Click 'Continue'. and Assign roles to the new users (e.g., Analyst, Admin).
   3. When done, click 'Send Invitation'.

After sending invitations, you can view pending users and their invite status in additional tabs.

{% hint style="info" %}
If an invitation expires, you can resend it from the pending users tab.
{% endhint %}

***

### 3. Managing User Roles

To update a user's role:

1. Click on the three dots menu next to the user's name.
2. Select 'Edit'.
3. Choose the new role you want to assign (e.g., Developer).
4. Click the confirmation button to save changes.

***

### 4. Deactivating Users

To deactivate or disable a user:

1. Click on the three dots menu next to the user's name.
2. Select 'Disable User'.

***

### 5. Auto-Join Email Domains

To allow users from specific email domains to join without an invitation:

1. Type the email domain in the designated field.
2. Click 'Add a Domain'.

To remove an auto-join domain:

1. Click on the three dots menu next to the domain.
2. Select 'Remove'.

{% hint style="info" %}
**Note:** Removing all auto-join domains means users will require an explicit invitation to join the workspace.
{% endhint %}

***

{% hint style="info" %}

#### Related Documentation

* [User Management](/app-help/documentation/settings/users)
  {% endhint %}

***

### Summary

You've now learned how to effectively manage users in your Paradime workspace, including inviting new users, managing roles, and setting up auto-join domains. These skills will help you maintain a secure and collaborative environment for your analytics engineering projects.

Next, we'll explore [Getting Started with the Paradime IDE](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide) to begin your journey with dbt™ development in Paradime.


# Getting Started with the Paradime IDE

The Paradime Integrated Development Environment (IDE) is a powerful, browser-based platform for real-time editing and execution of your dbt™️ project. With Paradime's IDE, you can seamlessly develop, run, test, and version control your dbt™️ project, while leveraging AI features to enhance your workflow.

**Estimated completion time:** 1.5 hours

***

### What you'll learn

<table data-view="cards"><thead><tr><th></th><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><p>🛠️ <strong>Setting Up a dbt™ Project</strong></p><p>Learn how to initialize a dbt™ project, connect a Git repository, and set up data warehouse connections in Paradime.<br></p></td><td><em>Estimated time: 10 minutes</em></td><td></td><td><a href="/pages/r9TExrNiHAIFrDhcQKNc">/pages/r9TExrNiHAIFrDhcQKNc</a></td></tr><tr><td><p>🏗️ <strong>Creating a dbt™ Model</strong></p><p>Create your first dbt™ model, from writing SQL to materializing it in your data warehouse.<br><br><em>Estimated time: 10 minutes</em></p></td><td></td><td></td><td><a href="/pages/xnzMpxvtdDSGzR3giyID">/pages/xnzMpxvtdDSGzR3giyID</a></td></tr><tr><td><p>🔍 <strong>Data Exploration in the Code IDE</strong></p><p>Discover how to use the Data Explorer and Scratchpad features for efficient data exploration and query iteration.<br><br>Estimated time: 10 minutes</p></td><td></td><td></td><td><a href="/pages/khbpNgFCkJqvE8HaAZwj">/pages/khbpNgFCkJqvE8HaAZwj</a></td></tr><tr><td>🚀 <strong>DinoAI: Accelerating Your Analytics Engineering Workflow</strong></td><td></td><td>Explore how DinoAI can supercharge your productivity in GitOps, data governance, and dbt™ development.<br><br><em>Estimated time: 45 minutes</em></td><td><a href="/pages/d2EMHuD6BFk1IyyEatul">/pages/d2EMHuD6BFk1IyyEatul</a></td></tr><tr><td>🔧 <strong>Utilizing Advanced Developer Features</strong></td><td></td><td>Learn about advanced tools that enhance data lineage, documentation, linting, and data handling in Paradime.<br><br>Estimated time: 45 minutes</td><td><a href="/pages/JAuamI7oHmds12eVYzn5">/pages/JAuamI7oHmds12eVYzn5</a></td></tr></tbody></table>

Let's get started with [Setting Up Your First dbt™ Project](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/setting-up-a-dbt-project).


# Setting Up a dbt™ Project

In this guide, you'll learn how to set up your development environment, initialize a dbt™ project, generate sources, and manage version control in Paradime. By the end of this tutorial, you'll have a fully functional dbt™ project ready for development.

**Estimated completion time:** 10 minutes

{% hint style="warning" %}

#### Prerequisites

* [Paradime developer or admin permissions](/app-help/documentation/settings/users/roles-and-permissions)
* [Access to a Git repository](/app-help/documentation/settings/git-repositories) (existing or new)
  {% endhint %}

### What you'll learn

In this guide, you'll learn how to:

1. [Connect your Git Repository](#id-1.-connect-your-git-repository)
2. [Initialize your dbt™ project](#id-2.-initialize-your-dbt-tm-project)
3. [Generate sources](#id-3.-generate-sources)
4. [Manage version control](#id-4.-version-control)

***

### 1. Connect your Git Repository

Before initializing your dbt™ project, ensure you have a Git repository connected to Paradime.

* **Existing dbt™ Repository**: If you have an existing dbt project in a git repository, make sure it's connected and proceed to the next section.
* **New Repository**: If you're starting from scratch, you'll create a new repository in the next step.

***

### 2. Initialize Your dbt™ Project

{% embed url="<https://youtu.be/GxDvYNw718w>" %}

To set up your dbt™ project, use the Paradime CLI (Terminal) by running the following command:

```
paradime repo init
```

This command will:

1. Create a branch called `initialize-dbt-project`.
2. Prompt you to name your dbt™ project (e.g., `demo_project`).
3. Set up a dbt project skeleton with all necessary folders and files, including `dbt_project.yml`.
4. Prompt you to generate `sources.yml` files.

{% hint style="info" %}
If you choose not to generate sources at this stage, you can do so later using DinoAI. See [DinoAI](/app-help/documentation/dino-ai) use case [*Creating Sources from your Warehouse*](/app-help/documentation/dino-ai/agent-mode/use-cases/creating-sources-from-your-warehouse) for step by step instructions.
{% endhint %}

***

### 3. Generate Sources

{% embed url="<https://youtu.be/42PtdhVpHyk>" %}

Generating `sources.yml` files is a crucial step in setting up your dbt™ project. These files define how dbt™ interacts with your data sources and ensure clear documentation, consistent data lineage, and robust error handling.

When prompted, Paradime will automatically fetch the available databases and schemas from your connected data warehouse. Complete the following steps:

1. **Select a database:** Navigate through the list using arrow keys. Highlight your source data database (e.g., `NBA`) and press `Enter`.
2. **Select one or more schemas:** In the schema list, use arrow keys to navigate. Press `>` to select one or more schemas containing your source tables (e.g., `Public`). Press `Enter` to confirm.

Once selected, Paradime will generate a `sources.yml` file in your project using the naming convention `sources_<schema_name>.yml`. You can rename these files to match your preferred naming convention.

***

### 4. Version Control

{% embed url="<https://youtu.be/8C0K7uTf79I>" %}

After making significant changes to your dbt™ project, it's important to keep your work synchronized with your main Git branch. Using Paradime's Git Lite feature, execute the following:

1. **Write Your Commit Message:** Use DinoAI's "Write Commit" feature to automatically generate a detailed commit message based on your specific code changes.
2. **Commit and Push:** Save your changes to your local repository, then push them to the remote repository to ensure your work is backed up and accessible.
3. **Open a Pull Request (PR):** Create a PR to allow your team to review and discuss the updates before merging.
4. **Merge the Pull Request:** Once the PR has been reviewed and approved, merge it into the main branch to finalize your changes.

***

{% hint style="info" %}

#### Related Documentation

* [Importing a repository](/app-help/documentation/settings/git-repositories/importing-a-repository)
* [Paradime CLI](/app-help/documentation/code-ide/terminal/paradime-cli)
* [Git Lite in Paradime](/app-help/documentation/code-ide/left-panel/git-lite)
* [DinoAI - Creating Sources from your Warehouse](/app-help/documentation/dino-ai/agent-mode/use-cases/creating-sources-from-your-warehouse)
  {% endhint %}

***

### Summary

You've successfully set up your dbt™ project in Paradime, including initializing the project, generating sources, and managing version control. Your project is now ready for development.

Next, we'll dive into [Creating a dbt™ Model](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/creating-a-dbt-model) to start building your data transformations.


# Creating a dbt™ Model

In this guide, you'll learn how to create your first dbt™ model in the Paradime IDE. By the end of this tutorial, you'll have a fully functional dbt™ model materialized in your data warehouse.

**Estimated completion time:** 10 minutes

{% hint style="info" %}

#### Prerequisites

* [Paradime developer or admin permissions](/app-help/documentation/settings/users/roles-and-permissions)
* [A dbt™ project set up in Paradime](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/setting-up-a-dbt-project)
* Basic understanding of SQL and dbt™ concepts
* Basic understanding of Git and version control concepts
  {% endhint %}

### What you'll learn:

* [Create a new file for your model](#id-2.-create-a-new-file)
* [Write a dbt™ model](#id-3.-write-a-dbt-tm-model)
* [Materialize your model](#id-5.-materialize-your-dbt-tm-model)

***

### 1. Create a New Branch

Before creating a model in your dbt™ project, it's essential to work on a new branch to keep your main branch clean and stable. Use Git Lite to create a new branch:

1. **Open Git Lite:** Click the source control icon within the left panel of the IDE.
2. **Create a New Branch:** Using the dropdown, select `+ New Branch` and give it a practical name (e.g., `my_first_dbt_model`).

### 2. Create a New File

{% embed url="<https://youtu.be/dMFjx85PK04>" %}

Create a new .sql file for your dbt™ model:

1. **Open your project file:** Click the folder Icon (📁) in the left panel to view your dbt™ project files.
2. **Create a new file:** Right-click on the folder where you want to add your new file (e.g., `models/sources`) and click `New File`.
3. **Name your file:** Use a descriptive name that reflects the purpose of the model (e.g., `nba_player_info.sql`).

### **3. Write a dbt™ Model**

{% embed url="<https://youtu.be/6Uvt7BCMMyg>" %}

With your new model open in the Code IDE, write SQL that transforms your data, using dbt's jinja syntax to reference sources and models dynamically. For example:

```sql
WITH source AS (
    SELECT 
        person_id AS player_id,
        first_name,
        last_name,
        team_name,
        position,
        height,
        weight
    FROM 
        {{ source('PUBLIC', 'COMMON_PLAYER_INFO') }}
    )
    
    SELECT 
        *
    FROM
        source
```

### 4. Materialize Your dbt™ Model

With your model written and previewed, it's time to materialize it in the data warehouse:

1. **Execute dbt run:** In the terminal at the bottom of your screen, run the command `dbt run`.
2. **Check for errors:** Review any errors or warnings that appear during the build process and resolve them as needed.

{% hint style="info" %}
You can also execute dbt commands from the commands panel. Click the "run model" dropdown and select your desired command for quick access to common dbt operations
{% endhint %}

### 5. Commit and Push

After successfully building your model, commit and push your work using the learnings from the [previous tutorial](https://docs.paradime.io/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/pages/r9TExrNiHAIFrDhcQKNc#id-4.-version-control).

1. **Write your commit message:** Use DinoAI's "Write Commit" feature to automatically generate a detailed commit message tailored to your specific code changes.
2. **Commit and Push:** Save your changes to your local repository, then push them to the remote repository to ensure your work is backed up and accessible.
3. **Open a Pull Request (PR):** Create a PR to allow your team to review and discuss the updates before merging.

***

{% hint style="info" %}

#### Related Documentation

* [Using Paradime's integrated terminal (CLI)](/app-help/documentation/code-ide/terminal)
* [Running dbt™ in Paradime's integrated terminal (CLI)](/app-help/documentation/code-ide/terminal/running-dbt)
  {% endhint %}

***

### Summary

Congratulations! You've successfully created your first dbt™ model in Paradime, from creating a new branch to committing your changes. Your model is now materialized in your data warehouse and ready for use.

Next, we'll explore [how to use the code IDE for data exploration](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/data-exploration-in-the-code-ide).


# Data Exploration in the Code IDE

The Paradime Code IDE offers powerful tools for data exploration and query iteration: the Data Explorer and Scratchpad. These features allow you to preview your data, test queries, and experiment with code efficiently.

**Estimated completion time:** 10 minutes

***

{% hint style="warning" %}

#### Prerequisites

* Basic understanding of SQL and dbt™ concepts
  {% endhint %}

***

### What You'll Learn

In this guide, you'll learn how to:

* Use the Data Explorer for compilation and data preview
* Utilize the Scratchpad for temporary code experiments
* Leverage these tools to improve your workflow efficiency

***

### **1. Data Explorer**

{% embed url="<https://youtu.be/Vr_Bb09u8Ag>" %}

The Data Explorer provides a Just in Time (JIT) live compiler, allowing you to preview compiled SQL without first compiling your entire dbt™️ project.

#### Key Features:

* **JIT Live Compiler**: Instantly see how your SQL compiles, including resolved macros and Jinja blocks.
* **Data Preview with Customizable Limits**: View query results directly in the IDE, with adjustable row limits.
* **CSV Download Option**: Easily export your preview data.
* **Support for Partial SQL Previews**: Test specific parts of your query by highlighting and previewing only the selected portion.

#### How to Use:

1. **Access the Data Explorer:** Click the 🔍 icon on the right panel to open the preview data view.
2. **View Compiled SQL:** Your compiled SQL will display automatically, showing resolved macros and Jinja blocks.

{% tabs %}
{% tab title="dbt™ model code" %}

```sql
-- This is the original dbt model code
-- It uses the ref() function to reference another model
SELECT
  *
FROM
  {{ ref('nba_player_info') }}
```

{% endtab %}

{% tab title="Compiled SQL  code" %}

```sql
-- This is the compiled SQL code
-- The ref() function has been resolved to the actual table name
SELECT
  *
FROM
  analytics.dbt_models.nba_player_info
```

{% endtab %}
{% endtabs %}

3. **Preview Data:** Click the "Preview Data" button to view your dbt™️ model or query results.
4. **Preview Partial SQL:** Highlight a specific portion of SQL and click "Preview Data" to view results for that portion.
5. **Adjust Query Limit:** Change the default limit (1-1000) in the "Query Limit" text box.
6. **Download Results:** Click "Download CSV" to save preview data.

***

### **2. Scratchpad**

{% embed url="<https://youtu.be/k0p560g1NRU>" %}

The Scratchpad feature allows for quick, temporary experiments in SQL and/or dbt™️, with instant data previews.

#### Key Features:

1. Temporary File Creation: Quickly create and experiment with code without affecting main project files.
2. Support for SQL and dbt™ Syntax: Write queries using both standard SQL and dbt™-specific functions.
3. Instant Data Preview: Leverage the Data Explorer for immediate results.
4. Session Persistence: Scratchpad files remain available across sessions.
5. Automatic Git Ignoring: Keep your repository clean by excluding scratchpad files from version control.

#### How to Use:

1. **Create a New Scratchpad File:**
   * Click the "New File" button at the top-right of your screen.
   * Paradime creates a file in the "paradime\_scratch" folder (e.g., "scratch-1", "scratch-2").
2. **Write and Test Code:**
   * Use SQL or dbt™️ syntax (including `ref()`, macros, etc.).
   * Leverage the Data Explorer for instant previews.
3. **Organize Your Scratchpad:**
   * Rename files as needed for better organization.
4. **Utilize Across Sessions:**
   * Scratchpad files remain available after logging out and back in.
5. **Maintain a Clean Git Repository:**
   * Scratchpad files are automatically gitignored.

***

{% hint style="info" %}

#### Related Documentation

* [Data Explorer](/app-help/documentation/code-ide/command-panel/data-explorer)
* [Scratchpad](/app-help/documentation/code-ide/additional-features/scratchpad)
  {% endhint %}

***

### Summary

You've learned how to use the Data Explorer for JIT compilation and data preview, and how to leverage the Scratchpad for temporary code experiments. These tools can significantly speed up your workflow, allowing you to quickly validate SQL queries and experiment with code without affecting your main project.

Next, we'll explore how to use [DinoAI for accelerating your analytics engineering workflow](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/dinoai-accelerating-your-analytics-engineering-workflow).


# DinoAI: Accelerating Your Analytics Engineering Workflow

Learn how DinoAI accelerates analytics engineering with two powerful modes. DinoAI's Agent Mode allows direct file editing.

As you advance in your dbt™ journey within Paradime's IDE, DinoAI is here to help you work smarter, not harder. DinoAI, Paradime's AI-powered assistant, is designed to enhance your analytics engineering workflow by automating key tasks, providing real-time support, and streamlining complex processes.

**Estimated completion time:** 45 minutes

{% hint style="info" %}
If you're brand new to Paradime, we recommend completing the earlier sections of the Paradime 101 Guide:

* [Getting Started with your Paradime Workspace](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide)
* [Setting Up a dbt™ Project](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/setting-up-a-dbt-project)
* [Creating a dbt™ Model](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/creating-a-dbt-model)
  {% endhint %}

{% embed url="<https://www.youtube.com/watch?v=DYqKm5XXO3I>" %}
High-level overview of DinoAI features
{% endembed %}

***

### What you'll learn

<table data-view="cards"><thead><tr><th></th><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td>🤖 <strong>DinoAI Agent</strong></td><td></td><td>Learn how to use DinoAI's Agent Mode to automate analytics engineering tasks with warehouse-aware AI that directly modifies your dbt™ files.<br><br><em>Estimated time: 15 minutes</em></td><td><a href="/pages/HogOc0dUCBi2wGk7Q1sQ">/pages/HogOc0dUCBi2wGk7Q1sQ</a></td></tr><tr><td>🛠️ <strong>Accelerating GitOps</strong></td><td></td><td>Boost your version control workflows with DinoAI's automated features, ensuring efficient and consistent Git operations.<br><br><em>Estimated time: 5 minutes</em></td><td><a href="/pages/OQxaVntzlJHDlrXp8BcE">/pages/OQxaVntzlJHDlrXp8BcE</a></td></tr><tr><td>📊 <strong>Accelerating Data Governance</strong></td><td></td><td>Enhance the quality, clarity, and consistency of your data assets by leveraging DinoAI's tools for automated testing, documentation, and visualization.<br><br><em>Estimated time: 10 minutes</em></td><td><a href="/pages/pg71WVRaOzuxz83w3AQ2">/pages/pg71WVRaOzuxz83w3AQ2</a></td></tr><tr><td>🚀 <strong>Accelerating dbt™ Development</strong></td><td></td><td>Speed up your dbt™ development process with DinoAI's capabilities, from model creation to debugging and SQL-to-dbt™ conversion.<br><br><em>Estimated time: 15 minutes</em></td><td><a href="/pages/CQjOPRCISeZNjzPMap7P">/pages/CQjOPRCISeZNjzPMap7P</a></td></tr></tbody></table>

***

### **Using DinoAI for General Inquiries**

In addition to accelerating specific tasks in your analytics engineering workflow, DinoAI's Ask Mode is your go-to assistant for virtually any inquiry, all without needing to leave the Code IDE. Whether you need quick answers or technical insights, DinoAI is ready to assist directly within your workflow.

You can ask DinoAI anything, including:

* "What time is it in Puerto Rico?"
* "How do I use the haversine function in Snowflake?"
* "How can I make this dbt model run more efficiently?"

With DinoAI, you get instant, context-aware support, so you can stay focused on your work without the distraction of context switching.

<figure><img src="/files/ahYoBOOfoO0VhFLQprgB" alt=""><figcaption></figcaption></figure>

Let's dive in!


# DinoAI Agent

The DinoAI Agent provides a flexible, conversational interface for working with DinoAI. It enables you to automate analytics engineering tasks with warehouse-aware AI.

The **DinoAI Agent** provides a flexible, conversational interface for working with DinoAI. It enables you to automate analytics engineering tasks with warehouse-aware AI that writes directly to your dbt™ project files.

### Why Use the DinoAI Agent?

Agent Mode transforms analytics engineering workflows by:

* **Eliminating Repetitive Tasks**: Automating up to 99% of rote work like column renaming, source updates, and documentation generation
* **Accelerating Development**: Creating models, tests, and documentation in seconds rather than hours
* **Ensuring Consistency**: Maintaining standards across your dbt™ project through DinoRules integration
* **Reducing Context Switching**: Eliminating the need to copy code between tools or query the warehouse separately

All this is achieved with **built-in guardrails** that prevent resource-intensive operations, protect your data warehouse, and ensure user oversight of all changes.

### How to Use Agent Mode

1. Click the **DinoAI icon** (🪄) in the right panel of the Code IDE.
2. Select "Agent Mode" at the bottom of the DinoAI panel.
3. Type your prompt.
4. Review and accept DinoAI's suggestion(s) to add or modify files directly in your project.

<figure><img src="/files/MvBj53xsMZ5aGanwDyY4" alt=""><figcaption></figcaption></figure>

### Key Features of Agent Mode

<table><thead><tr><th width="239.46875">Feature</th><th>Description</th></tr></thead><tbody><tr><td><strong>Direct File Editing</strong></td><td>Writes SQL models, YAML files, tests, and documentation directly into your dbt™ project without copy-pasting. Creates new files and updates existing ones seamlessly.</td></tr><tr><td><strong>Warehouse Awareness</strong></td><td>Understands your data warehouse structure by querying metadata from information schema. Identifies available tables and columns, recognizes data types, and works with multiple warehouse types (Snowflake, BigQuery, Redshift, Databricks, etc.).</td></tr><tr><td><strong>Add Context</strong></td><td>Allows you to explicitly target specific files or folders for DinoAI to consider. Provides precise control by letting you add specific files, active files, or entire directories as context.</td></tr><tr><td><strong>Built-in Guardrails</strong></td><td>Prevents resource-intensive operations with safety measures and requires user approval before creating, modifying, or running files.</td></tr><tr><td><strong>.dinorules Integration</strong></td><td>Honors your .<a href="/pages/T2LdXbm8NEruZg2TekIm">dinorules configuration file</a>, ensuring consistent formatting, naming conventions, and project structure across your entire codebase.</td></tr></tbody></table>

### Use Cases

Explore the following use cases to see Agent Mode in action:

* Creating DBT Sources from Data Warehouse
* Generating Well-Structured Base Models
* Building Complex Intermediate/Marts Models
* Bulk Documentation Generation
* Data Pipeline Configuration


# Creating dbt Sources from Data Warehouse

This guide demonstrates how DinoAI Agent streamlines the process of creating and updating sources.yml files by automatically extracting metadata from your data warehouse.

Setting up source definitions in dbt™ projects involves querying information schemas, documenting tables and columns, ensuring proper YAML formatting, and keeping sources up-to-date as new tables are added. This manual process can be time-consuming and prone to errors.

DinoAI Agent can automatically generate complete and accurate sources.yml files by directly accessing your data warehouse metadata.

{% embed url="<https://youtu.be/sAAAUo4OawM?si=7jVJPA2g48SnLhVX&t=26>" %}

### Example Prompt

> I uploaded some new data to my data warehouse . Can you create a sources.yaml file?

{% hint style="info" %}
**Optional**: When updating existing sources with new tables, you can add context by selecting your existing sources.yml file. Context helps DinoAI understand your existing project structure and make more relevant changes.
{% endhint %}

<figure><img src="/files/JQFbpLvDHIpVNKdalp3u" alt=""><figcaption></figcaption></figure>

### How It Works

After you enter your prompt:

1. DinoAI connects to your data warehouse and scans available schemas and tables
2. It retrieves column information including data types
3. If configured, DinoAI applies your .dinorules preferences
4. It generates a properly formatted sources.yml file

{% hint style="info" %}
**Note**: If you're updating existing (not creating) a YAML file, DinoAI preserves existing documentation and adds only the new tables
{% endhint %}

### Example Output

DinoAI will generate a properly formatted sources.yml file like this:

```yaml
version: 2

sources:
  - name: formula_one
    database: FORMULA_ONE_DB
    schema: raw
    tables:
      - name: CIRCUITS
        columns:
          - name: CIRCUITID
            data_type: NUMBER
          - name: CIRCUITREF
            data_type: VARCHAR
          # Additional columns...
            
      - name: CONSTRUCTORS
        columns:
          - name: CONSTRUCTORID
            data_type: NUMBER
          # Additional columns...
      
      # Additional tables...
```

### Key Benefits

* **Time Savings**: Reduces a 30+ minute manual task to seconds
* **Accuracy**: Eliminates typos and formatting errors
* **Maintainability**: Makes it easy to keep sources up-to-date as your warehouse evolves
* **Completeness**: Captures all tables and columns without missing anything

### When to Use This

* When setting up a new dbt™ project
* When data engineers have added new tables to your warehouse
* During data migrations or schema updates
* Any time your source data structure changes


# Generating Base Models

This guide shows how DinoAI Agent creates standardized base models from raw source data by applying consistent naming conventions and formatting rules.

Analytics engineers need to create base models with readable column names, consistent naming conventions, and appropriate transformations. This repetitive process can be tedious, especially with tables containing many columns.

DinoAI Agent can automatically generate well-structured base models that transform raw source data into more user-friendly formats.

{% embed url="<https://youtu.be/sAAAUo4OawM?si=nkLOSB_RSy9EEJNh&t=150>" %}

### Example Prompt

> I want you to create a set of base dbt models. These base models should rename columns that are confusing.

{% hint style="info" %}
**Optional**: You can add context by selecting your sources.yml file or existing base models. Context allows DinoAI to understand your naming conventions and maintain consistency with your existing files.
{% endhint %}

<figure><img src="/files/JQFbpLvDHIpVNKdalp3u" alt=""><figcaption></figcaption></figure>

### How It Works

After you enter your prompt:

1. DinoAI analyzes your sources.yml file and any provided context
2. It creates appropriate folder structure for base models
3. For each source table, it generates a model with readable column names
4. If configured, DinoAI will apply your .dinorules preferences

{% hint style="info" %}
**Note**: If you've already established naming patterns in your project, adding context helps DinoAI maintain those patterns in new models
{% endhint %}

### Example Output

DinoAI will generate base models like this:

```sql
WITH source AS (
    SELECT * FROM {{ source('formula_one', 'CIRCUITS') }}
)

SELECT
    circuitid AS circuit_id,
    circuitref AS circuit_reference,
    name AS circuit_name,
    location AS location,
    country AS country,
    lat AS latitude,
    lng AS longitude,
    alt AS altitude,
    url AS circuit_url
FROM source
```

### Key Benefits

* **Consistency**: Ensures all base models follow the same patterns and conventions
* **Readability**: Makes column names more user-friendly and easier to understand
* **Efficiency**: Creates dozens of models in seconds rather than hours
* **Standards**: Applies your team's SQL formatting and naming conventions automatically

### When to Use This

* When setting up initial base models for a new project
* After adding new source tables to your project
* When standardizing naming conventions across your project
* When needing to improve the readability of your raw data


# Building Intermediate/Marts Models

This guide demonstrates how DinoAI Agent helps create analytical models that join multiple tables with proper relationships and business logic.

Developing intermediate and marts models requires understanding table relationships, writing complex joins, applying business logic, and ensuring consistent formatting. This process demands careful consideration of data relationships and performance implications.

DinoAI Agent can analyze your existing models, understand their relationships, and create properly structured intermediate models with appropriate joins.

{% embed url="<https://youtu.be/vIiTrtbZ8cE?si=TayyUPIKejGUns9K>" %}

### Example Prompt

> I want to create an intermediate model which makes use of or allows me to understand data across races as well as constructor and driver standings for the season.

{% hint style="info" %}
**Optional**: Add context by selecting the relevant base models you want to join. Adding specific files as context helps DinoAI understand exactly which tables to include and how they should be related.
{% endhint %}

<figure><img src="/files/JQFbpLvDHIpVNKdalp3u" alt=""><figcaption></figcaption></figure>

After you enter your prompt:

1. DinoAI analyzes the structure of your provided models
2. It identifies the relationships between tables (e.g., foreign keys)
3. It creates a new model with properly structured CTEs and join conditions
4. If configured, DinoAI applies your .dinorules preferences

### Example Output

DinoAI will generate professional quality model(s). See example below:

<details>

<summary>Intermediate Model example</summary>

```sql
{{ config(materialized='table') }}

WITH races AS (
    SELECT * FROM {{ ref('stg_f1_races') }}
),

circuits AS (
    SELECT * FROM {{ ref('stg_f1_circuits') }}
),

results AS (
    SELECT * FROM {{ ref('stg_f1_results') }}
),

constructors AS (
    SELECT * FROM {{ ref('stg_f1_constructors') }}
),

driver_standings AS (
    SELECT * FROM {{ ref('stg_f1_driver_standings') }}
),

-- Join race data with circuit information
race_details AS (
    SELECT
        r.race_id,
        r.year,
        r.round,
        r.name AS race_name,
        c.circuit_name,
        c.location,
        c.country,
        r.date AS race_date
    FROM races r
    LEFT JOIN circuits c ON r.circuit_id = c.circuit_id
),

-- Join results with constructors
race_results AS (
    SELECT
        res.race_id,
        res.driver_id,
        res.constructor_id,
        con.constructor_name,
        con.nationality AS constructor_nationality,
        res.grid,
        res.position,
        res.points
    FROM results res
    LEFT JOIN constructors con ON res.constructor_id = con.constructor_id
)

-- Final combined query
SELECT
    rd.race_id,
    rd.year,
    rd.round,
    rd.race_name,
    rd.circuit_name,
    rd.country,
    rd.race_date,
    rr.constructor_id,
    rr.constructor_name,
    rr.constructor_nationality,
    rr.driver_id,
    rr.grid,
    rr.position,
    rr.points,
    ds.position AS driver_standing_position,
    ds.points AS driver_total_points
FROM race_details rd
JOIN race_results rr ON rd.race_id = rr.race_id
LEFT JOIN driver_standings ds ON rr.driver_id = ds.driver_id AND rd.race_id = ds.race_id
ORDER BY rd.year, rd.round, rr.position
```

</details>

### Key Benefits

* **Relationship Understanding**: Correctly identifies and implements join conditions
* **Code Organization**: Creates well-structured CTEs that make the logic easy to follow
* **Proper Formatting**: Maintains consistent SQL style according to your standards
* **Time Savings**: Reduces complex model development from hours to minutes
* **Visualization**: Can generate mermaid diagrams to visualize data flow when using .dinorules

### When to Use This

* When building analytical models that combine multiple data sources
* When implementing complex business logic across several tables
* When standardizing existing intermediate/marts models
* When exploring new analytical capabilities from your existing data


# Documentation Generation

This guide shows how DinoAI Agent automates the creation of documentation for multiple models simultaneously, saving hours of manual documentation work.

Documentation is often neglected in data projects because it's time-consuming to write descriptions for every model and column. Teams frequently postpone documentation or conduct separate "documentation sprints" that take days to complete.

DinoAI Agent can automatically generate comprehensive documentation for entire folders of models, complete with column descriptions and appropriate tests.

{% embed url="<https://youtu.be/koKhcpuKlx8?si=z9sj8K_yfEpBC7t5>" %}

### Example Prompt

> I have a bunch of new files in my Marts folder. Can you document this for me?

{% hint style="info" %}
**Optional**: Add context by selecting a directory containing the models you want to document. Using directory context is especially powerful for documentation tasks as it allows DinoAI to document multiple models at once.
{% endhint %}

<figure><img src="/files/JQFbpLvDHIpVNKdalp3u" alt=""><figcaption></figcaption></figure>

### How It Works

After you enter your prompt:

1. DinoAI scans all models in the specified directory
2. It analyzes each model's structure, column names, and relationships
3. It generates schema.yml files with model descriptions, column descriptions, and tests
4. If configured, DinoAI follows your .dinorules documentation standards

### Example Output

DinoAI will generate professional quality documentation (.yml files). See example below:

<details>

<summary>.yml file example</summary>

```yaml
version: 2

models:
  - name: int_f1_race_results_by_constructor
    description: "This model combines race results with constructor information to analyze performance by constructor across different races and seasons."
    columns:
      - name: race_id
        description: "Unique identifier for each race event"
        tests:
          - not_null
          
      - name: year
        description: "The year in which the race took place"
        tests:
          - not_null
          
      - name: race_name
        description: "The official name of the race event"
        
      # Additional columns with descriptions and tests...
```

</details>

### Key Benefits

* **Comprehensive Coverage**: Generates documentation for all models at once
* **Consistency**: Maintains a uniform documentation style across projects
* **Test Integration**: Automatically adds appropriate tests based on data types and relationships
* **Time Savings**: Turns days of documentation work into minutes
* **Adoption**: Makes it easy to keep documentation up-to-date as models evolve

### When to Use This

* Before sharing models with stakeholders
* When preparing for project handovers
* During documentation clean-up efforts
* After creating new models or making significant changes
* When implementing testing strategies


# Data Pipeline Configuration

This guide demonstrates how DinoAI Agent helps configure your dbt™ project settings to align with best practices and your folder structure.

Properly configuring dbt™ projects involves managing folder-specific materializations, schema naming conventions, and consistent tagging. As projects grow, maintaining these configurations becomes increasingly complex.

DinoAI Agent can analyze your project structure and automatically update your dbt\_project.yml file to implement best practices for materializations, schemas, and tags.

{% embed url="<https://youtu.be/koKhcpuKlx8?si=kYA2xQe4VXYwutVA&t=218>" %}

### Example Prompt

> Can you update my dbt\_project.yml to map all my models to my folder structure so that each SQL file in the respective folder will write to the respective schema and also make sure they are tagged?

{% hint style="info" %}
**Optional**: Add context by selecting your existing dbt\_project.yml file. This context is useful for configuration tasks as it allows DinoAI to understand your current settings before making changes.
{% endhint %}

<figure><img src="/files/JQFbpLvDHIpVNKdalp3u" alt=""><figcaption></figcaption></figure>

### How It Works

After you enter your prompt:

1. DinoAI analyzes your project's folder structure and existing configuration
2. It identifies model types based on folder organization (staging, intermediate, marts, etc.)
3. It generates appropriate configuration sections with schema naming and materializations
4. If configured, DinoAI applies your .dinorules preferences for project organization

{% hint style="info" %}
**Note**: Adding your existing dbt\_project.yml as context ensures DinoAI preserves your custom settings while adding new configurations
{% endhint %}

### Example Output

DinoAI will generate configuration like this:

```yaml
# dbt_project.yml update

models:
  formula_one:
    # Default materialization and schema settings
    +materialized: view
    +schema: "{{ target.schema }}"
    
    # Source models
    source:
      +schema: "{{ target.schema }}_source"
      +tags: ["source", "formula_one"]
    
    # Staging models
    staging:
      +schema: "{{ target.schema }}_staging"
      +tags: ["staging", "formula_one"]
      # Sub-folders within staging
      f1:
        +tags: ["f1"]
    
    # Intermediate models
    intermediate:
      +schema: "{{ target.schema }}_intermediate"
      +tags: ["intermediate", "formula_one"]
    
    # Marts models (typically materialized as tables)
    marts:
      +materialized: table
      +schema: "{{ target.schema }}_marts"
      +tags: ["marts", "formula_one", "reporting"]
      # Sub-folders within marts
      f1:
        +tags: ["f1"]
```

### Key Benefits

* **Optimized Performance**: Ensures efficient materialization strategies
* **Organizational Clarity**: Makes project structure more intuitive
* **Consistent Naming**: Implements standardized schema naming
* **Logical Tagging**: Makes it easier to run specific model groups
* **Maintainability**: Creates a configuration that scales with your project

### When to Use This

* When setting up a new dbt™ project
* When restructuring or reorganizing existing projects
* After adding new model categories or folders
* When implementing or updating tagging strategies
* Before optimization efforts to ensure proper materializations


# Using .dinorules to Tailor Your AI Experience

This guide explains how to use .dinorules files to customize DinoAI's behavior and ensure it follows your team's coding standards and best practices.

When working with dbt™ projects, teams often establish standards for SQL formatting, naming conventions, documentation, and modeling patterns. Ensuring all team members follow these standards consistently can be challenging.

.dinorules allow you to configure how DinoAI Agent operates within your project by establishing standards that it will automatically follow when generating or modifying code.

{% embed url="<https://youtu.be/HS1NKGqz7Y4?si=2mAG6br80Cw4S3PN&t=292>" %}

{% hint style="info" %}
**Benefits of using .dinorules**

* **Define Project-Specific Rules**: Customize DinoAI's behavior to match your team's unique needs and workflows.
* **Set Technical Standards**: Specify coding patterns, architecture guidelines, and best practices for your project.
* **Adapt Over Time**: Easily modify the rules as your project evolves and requirements change.
* **Team Collaboration**: Files are **git-tracked by default**, enabling shared standards across your entire data team.
  {% endhint %}

### How to Create a .dinorules File

Simply add a file named `.dinorules` to the root of your dbt™ project repository. The file uses natural language instructions and doesn't require any specific formatting, structure, or syntax. You can edit it just as you would a regular text document - use paragraphs, bullet points, or any organization that makes sense to you.

1. From the Code IDE, Click the DinoAI icon (🪄) icon on the left-side panel.\
   2\. Click the settings icon (⚙️) at the top right of the DinoAI panel. This automatically creates a new file named `.dinorules`.

{% hint style="warning" %}
Make sure the `.dinorules` file is placed in the **root directory of your repository**

```
your-repository
├── dbt_project/
│   ├── staging/
│   └── marts/
├── macros/
├── seeds/
├── .dinorules              # .dinorules file location
├── README.md
```

{% endhint %}

3. Add your [custom instructions to the file](https://docs.paradime.io/app-help/documentation/dino-ai-copilot/dbt-development/custom-rules-for-dinoai#example-.dinorules-file-configuration). These can be general or highly specific - there's no set syntax required.

{% @arcade/embed url="<https://app.arcade.software/share/9RN6tO67l5JPmEp1w8cm>" flowId="9RN6tO67l5JPmEp1w8cm" %}

#### Example DinoRules File

```
SQL Formatting
- Keyword capitalization (uppercase/lowercase)
- Comma style (leading/trailing)
- Indentation (spaces vs. tabs, amount)
- Aliasing conventions
- CTEs vs. subqueries preferences
- Comment requirements for complex logic

Naming Conventions
- Model naming patterns
- Column naming standards
- Test naming formats
- Documentation file naming

Documentation Standards
- Documentation file organization (single file vs. multiple files)
- Required column descriptions
- Standard test applications
- Metadata requirements

Modeling Patterns
- Dimensional modeling preferences
- Materialization defaults by folder
- Incremental model patterns
- Custom configurations

Visualization Standards
- When to include mermaid diagrams
- What elements to include in visualizations
- Visualization formatting
```


# Using Your Project as Context to Set Up .dinorules

This guide shows how to leverage your existing dbt™ project files to create comprehensive .dinorules that reflect your team's established patterns and conventions.

Creating effective [.dinorules](/app-help/documentation/dino-ai/dino-rules) from scratch can be challenging when you don't know what standards to define or how to articulate your team's existing patterns. Rather than starting with a blank file, you can analyze your current dbt™ project to identify established conventions and use those insights to build comprehensive rules.

DinoAI Agent can examine your existing models, YAML files, and project structure to automatically generate .dinorules that capture your team's actual practices.

{% embed url="<https://www.youtube.com/watch?v=UnugV57KKlE>" %}

#### Example Prompt

> Analyze my selected dbt™ project files and create a .dinorules that capture my existing SQL formatting, naming conventions, folder structure, and documentation patterns.

<figure><img src="/files/IYCgF5EceffI0rAG0xON" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
**Recommended**: Add [context](/app-help/documentation/dino-ai/context) by selecting multiple files that represent your best practices. include well-structured models from different folders (staging, marts), your dbt\_project.yml, and existing schema files. This gives DinoAI a comprehensive view of your standards.
{% endhint %}

### How to use

To generate a .dinorules file that matches your team standards, do the following:

1. **Open DinoAI** by clicking the DinoAI icon (🪄) in the right panel
2. **Enter prompt** - You can use the [above example](#example-prompt) or create your own
3. **Add files as context** using the "@" icon in the DinoAI panel. Select representative files from different folders.
   1. **Examples**: staging models, intermediate models, marts models, your dbt\_project.yml, and schema.yml files

### How It Works

After you enter your prompt:

1. DinoAI analyzes the structure and patterns in your selected files
2. It identifies consistent naming conventions, SQL formatting styles, and organizational patterns
3. It examines your folder structure and materialization strategies
4. It creates a [.dinorules](/app-help/documentation/dino-ai/dino-rules) file that codifies these discovered patterns
5. The generated rules ensure future AI-generated code matches your established standards

#### Example Output

DinoAI will generate a comprehensive .dinorules file like this:

<details>

<summary>Example .dinorules file</summary>

```
Project Context
- Domain: E-commerce analytics
- Data Stack: dbt™ with Snowflake
- Project Structure: source → staging → intermediate → marts hierarchy

SQL Formatting Standards (Based on Analyzed Files)
- Keywords: Always uppercase (SELECT, FROM, WHERE, etc.)
- Indentation: 4 spaces, no tabs
- Comma Style: Trailing commas in SELECT statements
- Aliasing: Always alias tables (use 'u' for users, 'o' for orders)
- CTEs: Prefer CTEs over subqueries for readability
- Comments: Add comments for complex business logic

Naming Conventions (Extracted from Current Models)
- Models: snake_case with prefixes (stg_, int_, dim_, fct_)
- Columns: snake_case, descriptive names
- Tables: Singular nouns (user, not users)
- Primary Keys: Always end with '_id' (user_id, order_id)

Folder Structure and Materializations
- staging/: Views only, basic transformations
- intermediate/: Views, complex business logic
- marts/: Tables or incremental, final business models
- All models tagged by domain and layer

Documentation Requirements
- All models require descriptions
- Business-critical columns need descriptions
- Primary keys need unique and not_null tests
- Foreign keys need relationships tests

Business Logic Patterns
- Always use UTC for timestamps
- Standardize null handling with COALESCE
- Use consistent date formatting (YYYY-MM-DD)
- Apply consistent rounding rules (2 decimal places for currency)

Quality Standards
- All models must compile without warnings
- Use consistent test patterns across similar model types
- Maintain referential integrity through proper testing
```

</details>

### Key Benefits

* **Consistency**: Ensures all future AI-generated code matches your existing standards
* **Discovery**: Reveals patterns you may not have consciously documented
* **Team Alignment**: Codifies implicit knowledge that experienced team members understand
* **Efficiency**: Creates comprehensive rules without starting from scratch
* **Accuracy**: Based on your actual codebase rather than generic best practices

### When to Use This

* When setting up .dinorules for the first time
* When your project has evolved organically and you want to formalize conventions
* Before major refactoring efforts to ensure consistency
* When preparing to scale your team and need documented standards


# AI-Generated Pull Request Descriptions

This guide shows how DinoAI Agent automatically generates comprehensive pull request descriptions by analyzing code changes, saving time and improving code review quality.

Writing meaningful pull request descriptions is often an afterthought for developers, leading to vague descriptions like "fix bug" or "update code" that provide no context for reviewers. This creates friction during code reviews and makes it difficult to understand the purpose and impact of changes.

DinoAI Agent can automatically generate detailed, professional pull request descriptions by analyzing the differences between your branch and main, creating comprehensive documentation that helps reviewers understand your changes.

{% embed url="<https://youtu.be/vf0rvfTgwaU?si=nzBLGpVjpAotiK7h&t=25>" %}

### How to Use

To generate a pull request description, use DinoAI's pre-built [.dinoprompt](/app-help/documentation/dino-ai/dino-prompts), *Generate a pull request description*.

1. **Open DinoAI** by clicking the DinoAI icon (🪄) in the right panel
2. **Access Prompt shortcut** by clicking the prompt button ("\[") within the DinoAI panel
3. **Create .dinoprompts file** (if needed): If you don't already have a .dinoprompts file, Paradime will offer to create one for you
4. **Select pre-built prompt** "Generate a pull request description"

<figure><img src="/files/hXdaAH6Hu83iia4rMuUl" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
**Note**: This prompt is pre-configured in your .dinoprompts file and can be updated per your team's standards. The prompt automatically uses `{{ git_diff_dinoai_test }}` to include all changes between your current branch and main.
{% endhint %}

### How It Works

1. **Automatic Diff Detection**: DinoAI automatically passes the diff between your current development branch and main branch as context ( `{{ git_diff_dinoai_test }}`- you don't need to manually specify what changes to analyze
2. **File Analysis**: DinoAI analyzes all modified files, added code, and removed code to understand the scope of changes
3. **Intelligent Generation**: DinoAI generates pull request descriptions that include summaries, detailed change lists, and relevant context for reviewers

### Example Output

DinoAI generates professional pull request descriptions like this:

```markdown
## Summary
Fixed invalid column identifier causing runtime errors in race analysis models and updated column naming for consistency.

## Changes Made
- Corrected `race_id` column reference in intermediate race results model
- Renamed `driver_name` to `driver_full_name` for clarity
- Updated `championship_position` to `championship_ranking_position`
- Modified `circuit_reference` field naming for consistency
- Resolved merge conflicts with concurrent changes

## Files Modified
- `models/intermediate/int_race_results.sql`
- `models/staging/stg_drivers.sql`
- `models/marts/mart_championship_standings.sql`

## Testing
- All models compile successfully
- Runtime errors resolved
- Column naming follows established conventions

## Impact
- Fixes production pipeline failures
- Improves data model readability
- Maintains consistency with existing naming patterns
```

### Key Benefits

* **Better Code Reviews**: Gives reviewers clear understanding of changes and their impact
* **Time Savings**: Eliminates the manual effort of writing detailed PR descriptions
* **Consistent Quality**: Maintains professional standards across all pull requests
* **Better Code Reviews**: Gives reviewers clear understanding of changes and their impact
* **Improved Searchability**: Detailed descriptions make it easier to find specific changes later

### When to Use This

* Before creating pull requests for any code changes
* When you need to document complex changes involving multiple files
* When working with team members who need context about your modifications
* For maintaining audit trails and change documentation
* Any time you want to improve the quality of your code review process


# Data Contract Checks with DinoAI & .dinorules

Use DinoAI and `.dinorules` to automatically enforce schema integrity whenever dbt™ code is generated or updated — before any file is touched.

{% hint style="info" %}
**What this does:** DinoAI runs a scoped column-level lineage scan on every structural change. The scan completes before DinoAI writes a single line of code, surfacing every breaking change as a structured impact report.

dbt™ model contracts catch type mismatches at compile time, but they can't detect a downstream model referencing a renamed column. DinoAI + `.dinorules` fills that gap automatically.
{% endhint %}

### How the scoped scan works

Instead of walking the full DAG, DinoAI uses a **column-level blast radius**: one level upstream to confirm the input schema is still compatible, and two levels downstream to catch the models most likely to break — but only if they actually *select* the changed column.

```
[ Source / seed ]  →  [ Changed model ]  →  [ +1 child ]  →  [ +2 child ]  →  [ Further downstream ]
     -1 upstream           trigger             scanned          scanned            NOT scanned ⚠
```

{% hint style="warning" %}
**Scoped by design:** the report always notes that only -1 / +2 was checked. If your change is in a widely shared staging model, manually verify anything beyond +2 before merging.
{% endhint %}

{% hint style="info" %}
**Column-aware only:** at each level DinoAI checks whether the downstream model's `SELECT` actually references the changed column. Models that `ref()` the changed model but don't use the column are skipped — no false positives.
{% endhint %}

***

### What triggers a check

A data contract check is automatically triggered whenever any of the following occur:

| Trigger                                                | Example                                                     |
| ------------------------------------------------------ | ----------------------------------------------------------- |
| Column renamed, removed, or type/logic changed         | `race_id` → `race_key` in `stg_f1__races.sql`               |
| Source schema modified                                 | Editing any `*_sources.yml` file                            |
| Staging or intermediate output columns altered         | Adding or removing a column from `stg_f1__account.sql`      |
| Mart grain, primary key, or key metric columns changed | Changing the grain of `fct_f1__constructor_race_result.sql` |
| dbt™ test added, removed, or modified on a key field   | Removing `not_null` from `constructor_id` in a `.yml` file  |
| Model deleted or renamed                               | Renaming `f1__account_master.sql`                           |

{% hint style="info" %}
The check is never skipped. It always runs before any file is generated or modified.
{% endhint %}

***

### What gets scanned

For every triggering change, the following are checked within the -1 / +2 window:

**-1 upstream**\
Confirm the direct parent model or source still exposes the expected column with a compatible type.

**+1 downstream (column-aware)**\
Find all direct children that `ref()` the changed model and check whether their `SELECT` uses the changed column. Skip those that don't.

**+2 downstream (column-aware)**\
Repeat the same column-aware check one level further. Only flag models where the column actually propagates.

**YAML documentation files**\
Verify that column entries in `_<model_name>.yml` still match the model's actual output at the changed layer.

**Tests**\
Confirm no tests on the changed column are orphaned within the scanned range.

{% hint style="info" %}
If the changed model references a seed within the -1 / +2 window, seed schema compatibility is also checked.
{% endhint %}

***

### Add the rule to .dinorules

In your project root, open or [create `.dinorules` ](/app-help/documentation/dino-ai/dino-rules)and add the following block:

```yaml
Data Contract Checks:
  # Triggers — runs before DinoAI writes any file
  - TRIGGER: Column renamed, removed, or type/logic changed in any SQL or YAML file
  - TRIGGER: Source schema modified in *_sources.yml
  - TRIGGER: Staging or intermediate output columns altered
  - TRIGGER: Mart grain, primary key, or key metric columns changed
  - TRIGGER: dbt test added, removed, or modified on a key field
  - TRIGGER: Model deleted or renamed

  # Scan scope — column-level lineage, -1 upstream and +2 downstream only
  - UPSTREAM: Check -1 parent model or source for input schema compatibility
  - DOWNSTREAM: Scan +1 and +2 children using column-level lineage only
      Only flag a downstream model if its SELECT actually references the changed column
      Skip models that ref() the changed model but do not use the column
  - YAML: Verify _<model_name>.yml column entries at the changed layer
  - TESTS: Confirm no tests on the changed column are orphaned within scan range
  - SEEDS: Check seed schema compatibility if referenced within the scan window

  # Reporting
  - Always produce a Data Contract Impact Report after the scan
  - Always include a scoped scan notice:
      "Scan limited to -1 / +2 column-level lineage.
       Verify models beyond +2 manually if this column is widely consumed."
  - If no assets are affected: "No downstream contract violations detected."

  # Enforcement
  - Never skip — complete the scan before generating or modifying any files
  - Prefer column aliasing over renaming for backward compatibility
  - Flag PK / grain changes on mart models as HIGH severity
  - Add ⚠️ Breaking Contract Changes section to PR description for breaking changes
```

{% hint style="info" %}
Commit `.dinorules` to version control alongside your dbt™ project. Every team member using DinoAI will automatically pick up the same contract enforcement rules.
{% endhint %}

***

### Reading the impact report

After every scan DinoAI produces a structured impact report. The format is always the same:

```markdown
## Data Contract Impact Report

### Triggering Change
- `wins` renamed to `total_race_wins` in models/staging/stg_f1__constructor_standings.sql

### Scan Scope
-1 upstream · +1 column-aware · +2 column-aware

### Affected Assets
| Level | Asset | File Path | Impact |
|-------|-------|-----------|--------|
| -1 | f1_raw.constructor_standings | models/staging/_stg_f1__sources.yml | Source column `wins` confirmed present ✓ |
| +1 | int_f1__race_results_standings | models/intermediate/int_f1__race_results_standings.sql | Selects `wins` — must update to `total_race_wins` |
| +2 | fct_f1__constructor_race_result | models/marts/fct_f1__constructor_race_result.sql | Does not select `wins` directly — no change needed ✓ |
| —  | _stg_f1__constructor_standings.yml | models/staging/_stg_f1__constructor_standings.yml | Column entry `wins` orphaned — update to `total_race_wins` |

### Action Required
- Update `int_f1__race_results_standings.sql`: replace `wins` with `total_race_wins`
- Update `_stg_f1__constructor_standings.yml`: rename column entry
- Re-run `dbt test -s int_f1__race_results_standings`

⚠ Scan limited to -1 / +2 column-level lineage.
  If `wins` is consumed beyond +2, verify those models manually before merging.
```

{% hint style="info" %}
If no downstream assets are affected, the report will explicitly state: **"No downstream contract violations detected."**
{% endhint %}

***

### Severity levels

| Severity      | When it applies                                                                          | Required action                                                                                    |
| ------------- | ---------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- |
| 🟢 **LOW**    | New column added — no downstream consumers affected within scan window                   | Document in PR description                                                                         |
| 🟡 **MEDIUM** | Non-key column renamed or removed — one or more +1 / +2 models select the changed column | Update affected consumers; prefer aliasing; include impact report in PR description                |
| 🔴 **HIGH**   | PK or mart grain changed                                                                 | Explicit reviewer sign-off + `⚠️ Breaking Contract Changes` section in PR + manual check beyond +2 |

{% hint style="warning" %}
For HIGH severity changes the scoped scan still applies — there is no automatic full-DAG walk. The scoped scan notice and reviewer sign-off requirement handle the residual risk for models beyond +2.
{% endhint %}

***

### Enforcement rules

| Rule                       | Detail                                                                                                                            |
| -------------------------- | --------------------------------------------------------------------------------------------------------------------------------- |
| **Never skip**             | The scan runs on every triggering event before any file is touched                                                                |
| **Column-aware only**      | At +1 and +2, only flag models whose `SELECT` actually uses the changed column; skip models that merely `ref()` the changed model |
| **Prefer aliasing**        | Keep the old column name as an alias while introducing the new one to preserve backward compatibility                             |
| **Always note scan scope** | The report must state that only -1 / +2 was checked and prompt manual verification further downstream                             |
| **HIGH severity sign-off** | PK / grain changes require explicit reviewer sign-off; scoped scan still applies                                                  |

### End-to-end workflow

The following example walks through renaming `wins` → `total_race_wins` in `stg_f1__constructor_standings.sql`.

{% stepper %}
{% step %}
**Make the change using an alias**

Keep the old column name as a backward-compat alias so +1 / +2 consumers keep working while the rename rolls out:

```sql
-- stg_f1__constructor_standings.sql
wins AS total_race_wins,
wins AS wins  -- backward-compat alias; remove after consumers are updated
```

{% endstep %}

{% step %}
**DinoAI detects the trigger and runs the scoped scan**

Before writing any file, DinoAI:

1. Checks `-1` for input schema compatibility
2. Walks `+1` column-aware — flags `int_f1__race_results_standings` because it selects `wins`
3. Walks `+2` column-aware — skips `fct_f1__constructor_race_result` because it does not select `wins` directly
   {% endstep %}

{% step %}
**Review the impact report and scope notice**

Check every flagged asset. Note the scoped scan warning at the bottom of the report. If the changed column is used in widely shared models, manually check anything beyond +2 before proceeding.
{% endstep %}

{% step %}
**DinoAI generates or updates the files**

With the scan complete, DinoAI writes the SQL and YAML changes. HIGH severity items pause generation and prompt for explicit confirmation before continuing.
{% endstep %}

{% step %}
**Ask DinoAI to run dbt™ tests and raise the PR**

Once the files are generated, ask DinoAI to run the tests and create the PR in one go:

```
Run dbt run and dbt test on all models affected by this change, then raise a PR.
Use the Data Contract Impact Report from this session as the content for the
## Data Contract Check section of the PR description.
```

Paste the impact report into the PR description under `## Data Contract Check`. Request at least one reviewer before merging.

DinoAI will:

1. Run `dbt run` and `dbt test` on all affected models and surface any failures before the PR is opened
2. Create the PR with the correct `[type]:` title based on the change made
3. Populate the PR description with the impact report — including the scan scope, affected assets table, action items, and the scoped scan notice — under `## Data Contract Check`
4. Request a reviewer if one is configured in your repo settings

{% hint style="warning" %}
For HIGH severity changes, DinoAI will add a `⚠️ Breaking Contract Changes` section to the PR description automatically and pause for your confirmation before opening the PR.
{% endhint %}
{% endstep %}
{% endstepper %}


# Keep docs in sync with model changes

DinoAI can automatically flag documentation gaps whenever you modify a dbt model — no more silently outdated YAML files.

### How it works

When you add this rule to your [`.dinorules` ](/app-help/documentation/dino-ai/dino-rules)file, DinoAI will:

1. Detect that a `.sql` file has been modified
2. Check whether the corresponding YAML doc file needs updating
3. Ask you before making any changes

### Setup

Add the following to your[ `.dinorules` ](/app-help/documentation/dino-ai/dino-rules)file:

```
Documentation Review on Model Changes:
    - Whenever a dbt model (.sql file) is modified, always check if the corresponding YAML documentation file (e.g. _<model_name>.yml) needs to be updated
    - Review whether any of the following have changed and may require doc updates:
      - Column additions, removals, or renames
      - Changes to business logic or transformations
      - Changes to model description or purpose
      - Changes to materialization or config
    - If documentation updates are needed, ask the user: "The model has changed — would you like me to update the documentation (descriptions, column definitions, etc.) to reflect these changes?"
    - Do not silently skip documentation review; always surface it as a step when modifying models
```

### What DinoAI will check

| Change type             | Example                  | Doc update likely needed? |
| ----------------------- | ------------------------ | ------------------------- |
| Column added or removed | New `revenue_usd` column | ✅ Yes                     |
| Column renamed          | `rev` → `revenue_usd`    | ✅ Yes                     |
| Business logic changed  | Filter added to a metric | ✅ Yes                     |
| Materialization changed | `view` → `table`         | ✅ Yes                     |
| Minor formatting fix    | Whitespace cleanup       | ❌ Probably not            |

### Example interaction

You modify `orders.sql` to add a new column. DinoAI will respond:

> "The model has changed — would you like me to update the documentation (descriptions, column definitions, etc.) to reflect these changes?"

Reply **yes** and DinoAI will update `_orders.yml` to match. Reply **no** to skip and move on.

### Why this matters

Without this rule, documentation drifts silently — columns get added, descriptions go stale, and the YAML file stops reflecting reality. Rule 15 makes doc review a first-class step in every model edit, not an afterthought.


# Accelerating GitOps

DinoAI's GitOps features help automate and streamline your version control workflows, ensuring your team stays on the same page with minimal effort.

**Estimated completion time:** 5 minutes

{% hint style="warning" %}

#### Prerequisites

* Basic understanding of Git and version control concepts
  {% endhint %}

***

### Using Dino AI to Generate Commit Messages

{% embed url="<https://www.youtube.com/watch?v=rq2TM1PRvSw>" %}

Leverage DinoAI to generate detailed, context-specific commit messages that enhance communication and clarity within your team. This feature is especially useful for teams working on collaborative projects where clear commit messages are crucial for tracking changes.

#### How to use:

1. **Make Changes to Your Code**: Begin by making the necessary changes in your project files.
2. **Click "Write Commit"**: In the Git Lite panel, select the "Write Commit" option to generate a commit message.
3. **Review and Edit**: Review the AI-generated commit message. If necessary, click "Write Commit" again for a new suggestion.
4. **Commit and Push**: Once satisfied, commit your changes and push them to your repository.

***

{% hint style="info" %}

#### Related Documentation

* [Working with Git Lite](/app-help/documentation/code-ide/left-panel/git-lite)
  {% endhint %}

***

### Summary

You've learned how to use DinoAI to generate detailed, context-specific commit messages for your Git workflows. This feature enhances team communication and improves version control history clarity. By leveraging DinoAI for commit messages, you can streamline your GitOps processes and maintain consistent, informative version control documentation.

Next, we'll explore [how DinoAI can accelerate your data governance practices](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/dinoai-accelerating-your-analytics-engineering-workflow/accelerating-data-governance).


# Accelerating Data Governance

Data governance is essential for maintaining the quality, clarity, and consistency of your data assets. DinoAI enhances your data governance practices by automating key tasks such as test generation, documentation, and entity relationship visualization.

**Estimated completion time:** 10 minutes

{% hint style="warning" %}

#### Prerequisites

* Basic understanding of dbt™ testing and documentation concepts
  {% endhint %}

### What you'll learn

In this guide, you'll learn how to use DinoAI for:

1. [Generating dbt™ Tests](#id-1.-generating-dbt-tm-tests)
2. [Autogenerating Data Documentation](#id-2.-autogenerating-data-documentation)
3. [Generating Entity Relationship Diagrams](#id-3.-generating-entity-relationship-diagrams)

***

### 1. Generating dbt™ Tests

{% embed url="<https://www.youtube.com/watch?v=G6bKF2Fffpk>" %}

Ensuring the reliability of your dbt™ models is a critical part of data governance. DinoAI simplifies this process by automatically generating comprehensive tests tailored to the structure and content of your dbt™ models. These tests help you maintain data integrity and consistency across your projects.

#### How to use:

1. **Open DinoAI**: Click the Dino AI icon (🪄) on the left side of the Editor.
2. **Access the Test Generation Feature**: Select the "One Click" command "Generate a test", or type "/test" in the prompt.

{% hint style="info" %}
If you use Elementary Data for dbt™ tests, try our elementary-specific command command, "Generate elementary tests for dbt model".
{% endhint %}

3. **Specify Your Model:** Enter the name of the model you want to test. For example:

`/Test @nba_player_info`

4. **Review Generated Test:** Carefully examine the AI-generated test.
5. **Implement Test:** Copy the generated test code and paste it into the appropriate .yml file in your project (e.g., `schema.yml`)
6. **Refine as Needed:** Edit and update the test as required for your specific use case.

{% hint style="success" %}
Alternative method to access the '/Test' command:

1. **Right-click** a .sql file in the project folder, files tab, or open file.
2. In the **DinoAI Copilot** dropdown, select "Generate tests".
   {% endhint %}

***

### 2. Autogenerating Data Documentation

{% embed url="<https://www.youtube.com/watch?v=0b4zjSF1gCw>" %}

Well-documented data assets are crucial for clarity and collaboration within any organization. DinoAI can automatically generate detailed documentation for your models and individual columns, saving time and ensuring consistency in your data documentation practices.

* **Open the Catalog Tab**: Navigate to the Catalog tab within the Apps Panel in the Code IDE.
* **Click "Autogenerate"**: At the top left of the Catalog panel, select the "Autogenerate" option to generate descriptions for your models and columns.
* **Edit and Save**: Review the generated descriptions, make any necessary edits to fit your project's specific context, and then save the changes.

***

### 3. Generating Entity Relationship Diagrams

{% embed url="<https://youtu.be/O0GoO0sOU5Q>" %}

Entity relationship diagrams (ERDs) are crucial for visualizing the structure and relationships within your data models. DinoAI can generate ERDs using Mermaid, allowing you to easily understand and communicate the relationships between different entities in your data.

#### How to use:

1. **Access the Mermaid Diagram Feature:**
   * Right-click a .sql file from the project folder, the files tabs, or within an opened .sql file.
   * Hover over the DinoAI Copilot dropdown and select "Generate Mermaid Diagram".
2. **Review Generated Code:** Carefully examine the AI-generated Mermaid code.
3. **Implement ER Diagram:** Copy the generated code and paste it into a new .mmd file in your project.
4. **Visualize**: Use a Mermaid viewer to see the visual representation of your model.

{% hint style="success" %}
Alternative method to access the Mermaid Diagram feature:

* **Open DinoAI:** Click the Dino AI icon (🪄) on the left side of the Editor.
* **Access the Mermaid Model Feature:** Select the "One Click" command "Generate a mermaid diagram for a dbt model", or type "/mermaid" in the prompt.
  {% endhint %}

***

### Related Documentation

* [DinoAI](broken://pages/kAxxwQWsHFubzexfR1N1)

***

### Summary

You've learned how to use DinoAI to generate dbt™ tests, autogenerate data documentation, and create entity relationship diagrams. These features significantly enhance your data governance practices, ensuring your data assets remain reliable, well-documented, and easy to understand.

Next, we'll explore [how DinoAI can accelerate your dbt™ development process](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/dinoai-accelerating-your-analytics-engineering-workflow/accelerating-dbt-development).


# Accelerating dbt™ Development

DinoAI enhances your dbt™ development by automating tasks such as model creation, explanation, debugging, and SQL-to-dbt™ conversion. This guide will help you leverage DinoAI to speed up your dbt™ development process, ensuring your models are well-structured, accurate, and aligned with your project standards.

**Estimated completion time:** 15 minutes

{% hint style="warning" %}

#### Prerequisites

* Basic understanding of dbt™ concepts and SQL
  {% endhint %}

***

### What you'll learn

In this guide, you'll learn how to use DinoAI for:

1. [Creating dbt™ Models](#creating-dbt-tm-models)
2. [Explaining dbt™ Models](#explaining-dbt-tm-models)
3. [Debugging dbt™ Models](#debugging-dbt-tm-models)
4. [Converting SQL to dbt™ Models](#converting-sql-to-dbt-tm-models)

***

### 1. Creating dbt™ Models

{% embed url="<https://www.youtube.com/watch?v=I_UXr2Rm1Wg>" %}

Creating dbt™ models is a foundational step in any dbt™ project. DinoAI simplifies this process by generating model code based on your prompts, ensuring consistency and saving time across your project.

#### How to use:

1. **Open DinoAI**: Click the Dino AI icon (🪄) on the left side of the Editor.
2. **Access the Create Model Feature**: Select the "One Click" command "Create a dbt model", or type "/model" in the prompt.
3. **Describe Your Model**: Enter a detailed prompt for the dbt model you'd like to create. For example:

*`/model Create a dbt model named int_nba_player_info that joins all columns from nba_player_info with the salary and season columns from nba_player_salaries, using the player_id column as the join key. Materialize it as a view.`*

4. **Review Generated Code:** Carefully examine the AI-generated model code.
5. **Implement the Model:** Copy the generated code and paste it into the appropriate .sql file in your project.
6. **Refine as Needed:** Modify the generated code to meet your specific requirements and project standards.

***

### 2. Explaining dbt™ Models

{% embed url="<https://www.youtube.com/watch?v=QRfL0P0A270>" %}

Understanding complex dbt™ models is essential for maintaining and collaborating on your data projects. DinoAI provides detailed explanations of your dbt™ models, breaking down their purpose, structure, and key components.

#### How to get started:

1. **Open DinoAI:** Click the Dino AI icon (🪄) on the left side of the Editor.
2. **Access the Explain Model Feature:** Select the "One Click" command "Explain a dbt model", or type "/explain" in the prompt.
3. **Specify Your Model:** Enter the name of the model you want explained. For example:

`/Explain nba_player_info`

4. **Review Explanation:** Carefully read the AI-generated summary of your model's purpose, output, and explanations of key parts like CTEs and subqueries.

{% hint style="success" %}
Alternative method to access the Copilot's '/Explain' command:

1. **Right-click** a .sql file in the project folder, files tab, or open file.
2. In the **DinoAI Copilot** dropdown, select "Explain model".
   {% endhint %}

***

### 3. Debugging dbt™ Models

{% embed url="<https://www.youtube.com/watch?v=PEhgGzEFlw4>" %}

Ensuring that your dbt™ models are free of errors is critical for the reliability of your data pipeline. DinoAI assists in debugging by identifying issues in your models and providing fixes.

#### How to use:

1. **Open DinoAI:** Click the Dino AI icon (🪄) on the left side of the Editor.
2. **Access the Debug Feature:** Select the "One Click" command "Debug a dbt model", or type "/fix" in the prompt.
3. **Specify Your Model:** Enter the name of the model you want to debug. For example:

`/Fix @nba_player_info`

4. **Review Changes:** Carefully examine the summary of changes made and the full, debugged code.
5. **Implement Fixes:** Copy the debugged code and paste it into your project's appropriate .sql file.
6. **Verify:** Use the Data Explorer to ensure the fixes work as expected.

{% hint style="success" %}
Alternative method to access the Copilot's '/Explain' command:

1. **Right-click** a .sql file in the project folder, files tab, or open file.
2. In the **DinoAI Copilot** dropdown, select "Fix model".
   {% endhint %}

***

### 4. Converting SQL to dbt™ Models

{% embed url="<https://www.youtube.com/watch?v=_OquYqJiMxA>" %}

Converting existing SQL queries into dbt™ models can save significant development time. DinoAI automates this conversion process, allowing you to quickly transition from raw queries to structured dbt™ models.

#### How to use:

1. **Right-click a .sql file:** You can right click a .sql file from the project folder, the files tabs, or within an opened .sql file.
2. **Hover over the DinoAI Copilot dropdown** and select option "Convert SQL to dbt model"
3. **Review Generated Code:** Carefully examine the AI-generated model code.
4. **Implement the Model:** Copy the generated code and paste it into the appropriate .sql file in your project.
5. **Refine as Needed:** Modify the generated code to meet your specific requirements and project standards.
6. **Verify:** Use the Data Explorer to ensure the model output is as expected.

{% hint style="success" %}
Alternative method to access the Copilot's '/sql\_to\_dbt' command:

1. **Right-click** a .sql file in the project folder, files tab, or open file.
2. In the **DinoAI Copilot** dropdown, select "Convert SQL to dbt model".
   {% endhint %}

***

### Summary

You've learned how to use DinoAI to accelerate your dbt™ development through model creation, explanation, debugging, and SQL conversion. These features speed up your workflow, enhance code quality, and improve collaboration, allowing you to focus more on data strategy.

Next, we'll learn about some [advanced Developer features in the Code IDE](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/utilizing-advanced-developer-features) to further improve your dbt™ development process.


# Utilizing Advanced Developer Features

Paradime offers advanced features to optimize your dbt™ development process. This section will guide you through tools that enhance data understanding, documentation, code quality, and data handling.

**Estimated completion time:** 45 minutes

{% hint style="info" %}
If you're brand new to Paradime, we recommend completing the earlier sections of the Paradime 101 Guide:

* [Getting Started with your Paradime Workspace](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide)
* [Getting Started with the Paradime IDE](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide)
  {% endhint %}

***

### What you'll learn

<table data-view="cards"><thead><tr><th></th><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><span data-gb-custom-inline data-tag="emoji" data-code="1f441">👁️</span> Visualize Data Lineage</td><td></td><td>Learn how to use Paradime's Data Lineage feature to understand the relationships between your dbt™ models.<br><br><em>Estimated time: 5 minutes</em></td><td><a href="/pages/p8c0D7DRoY7kNGb3Lhb7">/pages/p8c0D7DRoY7kNGb3Lhb7</a></td></tr><tr><td>📚 Autogenerate Data Documentation</td><td></td><td>Explore how to use Paradime's Catalog feature, powered by DinoAI, to effortlessly create and maintain high-quality documentation for your data assets.<br><br><em>Estimated time: 10 minutes</em></td><td><a href="/pages/KEucloZMijK7LEboKI6N">/pages/KEucloZMijK7LEboKI6N</a></td></tr><tr><td>🧹 Enforce SQL &#x26; YAML Best Practices</td><td></td><td>Learn how to use Paradime's integrated linting tools to maintain consistent, high-quality SQL and YAML code in your dbt™ project.<br><br><em>Estimated time: 20 minutes</em></td><td><a href="/pages/yZdgRthUJy8ggoLibSns">/pages/yZdgRthUJy8ggoLibSns</a></td></tr><tr><td>📁 Working with CSV Files</td><td></td><td>Discover how to effectively work with CSV files in Paradime using the integrated Rainbow CSV feature.<br><br><em>Estimated time: 10 minutes</em></td><td><a href="/pages/U6CTRYWOpXpATQxBZihl">/pages/U6CTRYWOpXpATQxBZihl</a></td></tr></tbody></table>

***

Let's dive into the first feature, [visualizing your data lineage](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/utilizing-advanced-developer-features/visualize-data-lineage)!


# Visualize Data Lineage

In this guide, you'll learn how to use Paradime's Data Lineage feature to understand the relationships between your dbt™ models. This visual representation is crucial for navigating and understanding your data pipeline, providing a clear view of the connections between your models and helping you assess the impact of changes before pushing to production.

**Estimated completion time:** 5 minutes

{% hint style="warning" %}

#### Prerequisites

* [A dbt™ project set up in Paradime](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/setting-up-a-dbt-project)
* [At least two dbt™ models materialized in your dbt™ project](https://docs.paradime.io/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/utilizing-advanced-developer-features/pages/xnzMpxvtdDSGzR3giyID#id-4.-materialize-your-dbt-tm-model)
  {% endhint %}

### What you'll learn

In this guide, you'll learn how to:

1. Access the Lineage View
2. Understand the Data Lineage visualization
3. Navigate through files using the Lineage View
4. Customize the Lineage View

***

{% embed url="<https://youtu.be/SeH5cLqcKJc>" %}

### Key Features

* **Immediate Dependencies View:** See upstream and downstream models at a glance.
* **File Navigation:** Quickly open related models directly from the lineage view.
* **Customizable Dependency Depth:** Adjust the number of visible upstream and downstream models.
* **Full Lineage Graph:** Expand to view the complete lineage of your project.

### How to Use

1. **Access the Lineage View**:
   * Open the Command Panel within the Code IDE (located at the bottom of the screen)
   * Click on the "Lineage" tab.
2. **Understand the Data Lineage:**
   * By default, the model currently open in your editor is at the center.
   * Upstream model dependencies are on the left.
   * Downstream model dependencies are on the right.
3. **Navigate Through Files:**
   * Hover over a node and click the 👁️ icon to open that model in the editor.
4. **Adjust Visible Dependencies:**
   * Use the text fields on the right to change the number of visible upstream and downstream dependencies.
   * Enter a number and press Enter, or use "+" for all dependencies.
5. **Customize the View:**
   * Use the zoom in/out and recenter icons on the right (or pinch your trackpad)
   * Click the expand button for a full-screen view.

By leveraging Data Lineage, you can quickly understand model dependencies, navigate your project efficiently, and assess the impact of changes before pushing to production.

***

{% hint style="info" %}

#### Related Documentation

* [Data Lineage](/app-help/documentation/code-ide/command-panel/lineage-preview)
  {% endhint %}

***

### Summary

You've learned how to use Paradime's Data Lineage feature to visualize and understand the relationships between your dbt™ models. By leveraging this tool, you can quickly understand model dependencies, navigate your project efficiently, and assess the impact of changes before pushing to production.

Next, we'll explore [how to autogenerate data documentation](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/utilizing-advanced-developer-features/auto-generated-data-documentation) to further enhance your dbt™ project management.


# Auto-generated Data Documentation

In this guide, you'll learn how to use Paradime's Catalog feature, powered by DinoAI, to effortlessly create and maintain high-quality documentation for your data assets. Well-documented assets are crucial for clarity and collaboration, and the Catalog feature helps you comprehensively document your dbt™ models and their columns, ensuring consistency across your project.

**Estimated completion time:** 10 minutes

{% hint style="warning" %}

#### Prerequisites

* [A dbt™ project set up in Paradime](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/setting-up-a-dbt-project)
* [At least one dbt™ model materialized in your dbt™ project](https://docs.paradime.io/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/utilizing-advanced-developer-features/pages/xnzMpxvtdDSGzR3giyID#id-4.-materialize-your-dbt-tm-model)
  {% endhint %}

### What you'll learn

In this guide, you'll learn how to:

1. Access and navigate the Catalog feature
2. View and edit model documentation
3. Use DinoAI for auto-generating documentation
4. Review and customize generated content

***

### Tutorial

{% embed url="<https://youtu.be/31RXmmxgcko>" %}

### Key Features

* **Model Classification**: Tag and categorize your models for easy organization.
* **Model Description**: Add detailed descriptions to explain the purpose and logic of each model.
* **Model Column Details**: Document metadata for each column, including type, tests, and classification.
* **DinoAI Auto-generation**: Automatically generate context-specific descriptions for models and columns.
* **Auto-updating YAML files**: All documentation updated via the Catalog UI will immediately be reflected in your project's corresponding .yaml files.

### How to Use

1. **Access the Catalog:**
   * Open the Command Panel within the Code IDE (located at the bottom of the screen)
   * Click on the "Catalog" tab.
2. **View and Edit Model Information:**
   * Click on any model in your project to view it's documentation.
   * Option: Click "Expand" icon on the top right of Catalog tab to expand view.
   * Click "Edit" Icon  (**✎)** to edit any model classifications, description, column details, etc.
3. **Use DinoAI for Auto-generation:**
   * In the Catalog tab, click "Autogenerate" at the top left.
   * DinoAI will generate a detailed, context-specific model model and column descriptions.
   * Click "Edit" Icon  (**✎)** to edit AI generated descriptions.
4. **Review and Customize Generated Content:**
   * Make any necessary edits to tailor the documentation to your specific needs.
   * Save your changes.

{% hint style="success" %}

#### Pro Tips

* Use auto-generation as a starting point, then customize for your specific needs.
* Regularly update documentation as your models evolve.
* Consistency in documentation style improves team collaboration.
  {% endhint %}

***

{% hint style="info" %}

#### Related Documentation

* [How to use the data catalog](/app-help/documentation/data-catalog)
* [Utilizing data assets](/app-help/documentation/data-catalog)
  {% endhint %}

***

### Summary

You've learned how to use Paradime's Catalog feature to autogenerate and maintain high-quality documentation for your dbt™ models. By leveraging this feature, you can significantly speed up your documentation process, ensure consistency across your project, and maintain up-to-date documentation as your models evolve.

Next, we'll explore [how to enforce SQL and YAML best practices](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/utilizing-advanced-developer-features/enforce-sql-and-yaml-best-practices) to further improve your code quality.


# Enforce SQL and YAML Best Practices

Paradime offers integrated tools like SQLFluff and Prettier to help enforce best practices for SQL and YAML files, ensuring high-quality code and reducing errors in your dbt™ project.\
\
**Estimated completion time:** 20 minutes

{% hint style="warning" %}

#### Prerequisites

* [A dbt™ project set up in Paradime](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/setting-up-a-dbt-project)
* [At least one .sql and .yml file in your dbt™ project](https://docs.paradime.io/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/utilizing-advanced-developer-features/pages/xnzMpxvtdDSGzR3giyID#id-4.-materialize-your-dbt-tm-model)
* Basic understanding of SQL and YAML syntax
  {% endhint %}

### What You'll Learn

In this guide, you'll learn how to:

* Use SQLFluff to lint and format SQL files
* Utilize Prettier to format and debug your YAML files
* Customize settings for both tools to match your team's standards

***

### 1. SQLFluff

SQLFluff is an integrated linting tool in Paradime that helps maintain consistent, high-quality SQL code. The [pre-configured template](/app-help/documentation/integrations/code-ide/sql-fluff#set-your-configuration-file-by-adding-supported-rules) is for Snowflake and dbt, setting basic rules for SQL formatting such as line length, indentation, aliasing, and capitalization. It provides a foundation for consistent SQL styling that can be easily customized to fit specific project needs.

{% embed url="<https://youtu.be/RVCTo8fX-7Q>" %}

**Key Features:**

* **Integrated Linting**: Automatically check SQL code against standard or custom rules.
* **Pre-configured for dbt™**: Comes ready to use with basic rules tailored for dbt™ projects.
* **Real-time Formatting**: Use the 'Prettier' button in Paradime's IDE for instant code corrections.

**How to Use SQLFluff:**

1. Select a .sql file within your dbt™ project.
2. Click the `Prettier` button in the commands panel to automatically format your .yml file.

<figure><img src="/files/OxKBuVbFsSi1fod6ahh5" alt=""><figcaption></figcaption></figure>

3. **Optional:** Customize your SQL formatting by creating a .sqlfluff file in your dbt™ root directory (this is in the same directory where your dbt\_project.yml lives).

<figure><img src="/files/oaBRoQf4I9fyKoHo0bMz" alt=""><figcaption></figcaption></figure>

{% hint style="warning" %}
If you don't have a `.sqlfluff` file in your project, simply create new file with the exact name ".sqlfluff" in the same directory where your `dbt_project.yml` lives.
{% endhint %}

#### Example SQLFluff Formatting

{% tabs %}
{% tab title="Before Formatting" %}

```sql
WITH player_info AS (SELECT * FROM {{ ref('nba_player_info') }})
, player_salaries AS (SELECT player_id, salary, season FROM {{ ref('nba_player_salaries') }})
, joined AS (SELECT pi.*, ps.salary, ps.season FROM player_info AS pi LEFT JOIN player_salaries AS ps ON pi.player_id = ps.player_id)
SELECT * FROM joined
```

{% endtab %}

{% tab title="After Formatting with SQLFluff:" %}

```sql
with player_info as (
    select *
    from {{ ref('nba_player_info') }}
),
player_salaries as (
    select
        player_id,
        salary,
        season
    from {{ ref('nba_player_salaries') }}
),
joined as (
    select
        pi.*,
        ps.salary,
        ps.season
    from player_info as pi
    left join player_salaries as ps on pi.player_id = ps.player_id
)

select *
from joined
```

{% endtab %}
{% endtabs %}

***

### 2. Prettier

Prettier is an integrated code formatter in Paradime that helps maintain consistent, error-free YAML files in your dbt™ project. Prettier comes pre-installed in your project and uses default configurations provided by the [Prettier library](https://prettier.io/docs/en/configuration.html), which can be easily customized to fit specific YAML preferences.

{% embed url="<https://youtu.be/vqh9M6rYB7A>" %}

**Key Features:**

* **Automatic Formatting**: Formats YAML files for improved readability and code quality.
* **Error Detection**: Highlights severe formatting errors and assists in debugging.
* **Customization**: Allows custom configurations through a .prettierrc file.

**How to Use:**

1. Select a .yml file within your dbt™ project.
2. Click the `Prettier` button in the commands panel to automatically format your .yml file.

<figure><img src="/files/OxKBuVbFsSi1fod6ahh5" alt=""><figcaption></figcaption></figure>

3. If more severe errors are detected, click the `Prettier` button in the toolbar to debug.

<figure><img src="/files/C2bALCYnpp7ImGXqHxm7" alt=""><figcaption></figcaption></figure>

4. **Optional**: Customize your YAML formatting by creating a .prettierrc file in your dbt™ root directory (this is in the same directory where your dbt\_project.yml lives).

<figure><img src="/files/22GGNWGaDjaX2PbAMVfJ" alt=""><figcaption></figcaption></figure>

#### Example YAML Formatting

{% tabs %}
{% tab title="Before formatting" %}

```yaml
version: 2

models:
  - name: nba_player_info
    columns:
      - name: player_id
        tests:
          -   unique
          - not_null
      - name: first_name
      - name: last_name
      - name: team_name
      - name: position

      - name:   height
      - name: weight
  - name: nba_player_salaries
    columns:
      - name: player_id
      - name: player_name
      - name: salary
      - name:   season

```

{% endtab %}

{% tab title="After Applying Prettier" %}

```yaml
version: 2

models:
  - name: nba_player_info
    columns:
      - name: player_id
        tests:
          - unique
          - not_null
      - name: first_name
      - name: last_name
      - name: team_name
      - name: position
      - name: height
      - name: weight
  - name: nba_player_salaries
    columns:
      - name: player_id
      - name: player_name
      - name: salary
      - name: season
```

{% endtab %}
{% endtabs %}

***

{% hint style="info" %}

#### Related Documentation

* [SQL Fluff](/app-help/documentation/integrations/code-ide/sql-fluff)
* [Code Quality](/app-help/documentation/code-ide/command-panel/code-quality)
* [Prettier](/app-help/documentation/integrations/code-ide/prettier)
  {% endhint %}

***

### Summary

By using SQLFluff and Prettier in Paradime, you enforce best practices across your SQL and YAML files. These tools help maintain consistent, clean, and high-quality code, improving your project's readability and reducing errors. Customize them to fit your team's coding standards, streamline your workflow, and keep your dbt™ project error-free.

Next, we'll explore [how to work with CSV files in Paradime](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/utilizing-advanced-developer-features/working-with-csv-files).


# Working with CSV Files

In this guide, you'll learn how to effectively work with CSV files in Paradime using the integrated Rainbow CSV feature. This tool enhances your ability to view, edit, and analyze CSV data directly within the Paradime IDE.

**Estimated completion time:** 10 minutes

{% hint style="warning" %}

#### Prerequisites

* [A dbt™ project set up in Paradime](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/setting-up-a-dbt-project)
* A .csv file to import into your dbt™ project
  {% endhint %}

### What You'll Learn

In this guide, you'll learn how to:

1. Access Rainbow CSV features
2. Align CSV columns
3. Perform multi-cursor editing
4. Set header lines
5. Customize separators
6. Use CSVLint for data quality checks

***

### Tutorial

{% embed url="<https://youtu.be/QwyjmL2BT2Q>" %}

### Key Features

* Automatic Color Coordination: CSV columns are color-coded for easy reading.
* CSV Column Alignment: Visually align columns for better organization.
* Multi-cursor Column Editing: Efficiently make bulk edits across columns.
* Header Line Freezing: Keep column names visible while scrolling.
* Customizable Separators: Choose how columns should be distinguished.
* CSVLint: Check for consistency in quotes usage and field count.

### How to Use

#### 1. Launch Rainbow CSV (H3)

1. If you don't already have a .csv file in your project, drag one into your project's repository.
2. Open a CSV file in the Paradime IDE.
3. In the status bar at the bottom of the IDE, turn on Rainbow CSV.

<figure><img src="/files/0qBvZpxXId3HOpBVztTo" alt=""><figcaption></figcaption></figure>

4. Open the Command Panel:

* On Mac: Press `⌘` + `Shift` + `P`
* On Windows: Press `Ctrl` + `Shift` + `P`

#### 2. Access the Commands Panel

To access all Rainbow CSV commands:

1. Open the Command Panel
2. Type "Rainbow CSV"
3. You'll see a list of available commands:

<figure><img src="/files/LITdfoQ2jAWHNZyTvG9p" alt=""><figcaption></figcaption></figure>

#### 3. Utilize Rainbow CSV Commands

Here are some key Rainbow CSV commands you can use:\ <br>

| Command                      | Description                                                                            | Use case                                                    | Where to Find it                 |
| ---------------------------- | -------------------------------------------------------------------------------------- | ----------------------------------------------------------- | -------------------------------- |
| **Set Header Line**          | Designates the current line as the header                                              | Improves column identification while scrolling              | Command Panel                    |
| **Align CSV Columns**        | Visually aligns columns without modifying the original file                            | Enhances readability, making it easier to analyze data      | Command Panel or Status bar (UI) |
| **Set Rainbow Separator**    | Allows selection of predefined separators (comma, tab, semicolon, pipe) or custom ones | Adapts to various CSV formats, improving column distinction | Command Panel                    |
| **Edit Column Before/After** | Enables multi-cursor column editing                                                    | Facilitates efficient bulk edits in specific columns        | Command Panel                    |
| **CSVLint**                  | Automatically checks for consistency in quotes usage and field count                   | Helps maintain data integrity and spot formatting errors    | Status Bar (UI)                  |

Discover all Rainbow CSV features via the command panel:

{% hint style="success" %}

#### Pro Tips

* Use Rainbow CSV for quick data quality checks before importing into your dbt™ models.
* Leverage multi-cursor editing for rapid data cleaning tasks.
* Combine Rainbow CSV with dbt™ seeds for efficient data loading and version control.
* Use CSVLint regularly to catch formatting issues early in your data pipeline.
  {% endhint %}

***

{% hint style="info" %}

#### Related Documentation

* [Rainbow CSV](/app-help/documentation/integrations/code-ide/rainbow-csv)
  {% endhint %}

***

### Summary

You've learned how to effectively work with CSV files in Paradime using the Rainbow CSV feature. This tool enhances your ability to view, edit, and analyze CSV data directly within the Paradime IDE, improving your data handling capabilities and workflow efficiency.\
\
Next, we'll explore[ how to run dbt™ in production with Bolt](/app-help/guides/paradime-101/running-dbt-in-production-with-bolt).


# Managing dbt™ Schedules with Bolt

Bolt is Paradime's powerful tool for running dbt™ in production environments. It allows you to efficiently schedule and manage your dbt™ jobs, ensuring your data transformations are executed reliably and on time.

**Estimated completion time:** 60 minutes

{% hint style="info" %}
If you're brand new to Paradime, we recommend completing the earlier sections of the Paradime 101 Guide:

* [Getting Started with your Paradime Workspace](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide)
* [Getting Started with the Paradime IDE](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide)
  {% endhint %}

***

### What you'll learn

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>🛠️ Creating Bolt Schedules</strong></td><td><p>Learn how to set up and configure dbt™ job schedules using the Bolt UI, including understanding different schedule types and triggers.<br></p><p>Estimated time: 10 minutes</p></td><td><a href="/pages/JLsspFwTVjUyE9px1ACE">/pages/JLsspFwTVjUyE9px1ACE</a></td></tr><tr><td>🔍 <strong>Understanding schedule types and statuses</strong></td><td><p>Explore the different types of schedules available (Standard, Deferred, and Turbo CI) and various ways to trigger your schedules.<br></p><p>Estimated time: 10 minutes</p></td><td><a href="/pages/C1Orgsnh1mRVDVaaHp0G">/pages/C1Orgsnh1mRVDVaaHp0G</a></td></tr><tr><td>📊 <strong>Viewing run history and analytics</strong></td><td><p>Discover how to monitor your dbt™ schedules, view run logs, and explore detailed analytics for your schedules.<br></p><p>Estimated time: 15 minutes</p></td><td><a href="/pages/awJgmDFlnMqzyoI34hrJ">/pages/awJgmDFlnMqzyoI34hrJ</a></td></tr><tr><td>🔔 <strong>Setting up notifications</strong></td><td><p>Learn how to configure Slack and email notifications to stay informed about your Bolt Schedule statuses.<br></p><p>Estimated time: 10 minutes</p></td><td><a href="/pages/F8g9TA0ECqBMQDZHnQoh">/pages/F8g9TA0ECqBMQDZHnQoh</a></td></tr><tr><td><strong>🐛 Debugging Failed Runs</strong></td><td>Learn how to identify, debug, and resolve failed dbt™ runs using Bolt's comprehensive logging and debugging tools.<br><br><em>Estimated time: 15 minutes</em></td><td><a href="/pages/kORvYWlNxCHOqUcgP9i9">/pages/kORvYWlNxCHOqUcgP9i9</a></td></tr></tbody></table>

Let's dive in and [create a bolt schedule](/app-help/guides/paradime-101/running-dbt-in-production-with-bolt/creating-bolt-schedules).


# Creating Bolt Schedules

Bolt provides an intuitive UI-based method to create and manage your dbt™ job schedules. This guide will walk you through the process of setting up a schedule using the Bolt UI.

{% hint style="info" %}
While this guide focuses on UI-based scheduling, Bolt also supports YAML-based scheduling for users who prefer configuration-as-code. For YAML-based scheduling, please refer to our [documentation](/app-help/documentation/bolt/creating-schedules#yaml-based-schedules).
{% endhint %}

**Estimated completion time:** 10 minutes

{% hint style="warning" %}

#### Prerequisites

* Access to the [Bolt application](https://www.paradime.io/pricing) in Paradime
* A [Schedule Connection](/app-help/guides/paradime-101/getting-started-with-your-paradime-workspace/setting-up-data-warehouse-connections) to your data warehouse (AKA Production Connection)
* [At least one dbt™ model materialized in your dbt™ project](https://docs.paradime.io/app-help/guides/paradime-101/running-dbt-in-production-with-bolt/pages/xnzMpxvtdDSGzR3giyID#id-4.-materialize-your-dbt-tm-model)
* Basic understanding of dbt™ commands
  {% endhint %}

### What You'll Learn

In this guide, you'll learn how to:

1. [Create a new Bolt schedule](#id-1.-creating-a-new-schedule)
2. [Configuring Bolt Schedules](#id-2.-configuring-your-schedule)

***

### 1. Creating a New Schedule

To create a new schedule, follow these simple steps:

1. Navigate to the Bolt application from the Paradime Home Screen.
2. Click on "+ New Schedule" and then "+ Create New Schedule"

{% @arcade/embed url="<https://app.arcade.software/share/Pi4oJO8JPzUzrBxhZNuT>" flowId="Pi4oJO8JPzUzrBxhZNuT" %}

***

### 2. Configuring Your Schedule

{% hint style="info" %}
This guide covers "Standard" schedules. We'll explore more advanced types like "Deferred" or "Turbo-CI" in the [next section](/app-help/guides/paradime-101/running-dbt-in-production-with-bolt/understanding-schedule-types-and-triggers).
{% endhint %}

{% embed url="<https://youtu.be/XlYaJehw0dg>" %}

Follow these steps to configure your schedule:

1. Fill out the required [schedule fields](#ui-based-schedule-fields)
2. Optionally, configure additional fields as needed
3. Click the `Save` button to publish your schedule

#### Schedule Fields <a href="#ui-based-schedule-fields" id="ui-based-schedule-fields"></a>

| Term                     | Description                                | Example                                                   | Required |
| ------------------------ | ------------------------------------------ | --------------------------------------------------------- | -------- |
| Type                     | Execution type for your scheduled dbt™ run | Standard, Deferred, or Turbo CI                           | Yes      |
| Name                     | Identifier for your dbt™ schedule          | hourly\_schedule                                          | Yes      |
| Commands                 | dbt™ commands to execute (can be multiple) | `dbt run`, `dbt test`                                     | Yes      |
| Git Branch               | Branch used for schedule execution         | main                                                      | Yes      |
| Owner Email              | Schedule owner's email                     | <me@email.com>                                            | Yes      |
| Trigger Type             | How the schedule is initiated              | Scheduled run, On Run Completion, On Merge, Cron Schedule | Yes      |
| Cron Schedule            | Frequency of schedule runs (UTC-based)     | @hourly                                                   | No       |
| Slack Notify On          | When to send Slack notifications           | failed and/or passed                                      | No       |
| Slack Channels and Users | Recipients of Slack alerts                 | #data-team-alerts                                         | No       |
| Email Notify On          | When to send email notifications           | failed and/or passed                                      | No       |
| Email Notify             | Email recipients for alerts                | <example@email.com>                                       | No       |

***

{% hint style="info" %}

#### Related Documentation

* [Understanding Schedule Types and Triggers](/app-help/documentation/bolt/managing-schedules/schedule-configurations#advanced-configuration-options)
* [YAML-based Scheduling](/app-help/documentation/bolt/creating-schedules#yaml-based-schedules)
  {% endhint %}

***

### Summary

You've learned how to create a new Bolt schedule, configure its settings, and understand the purpose of different schedule fields. This knowledge forms the foundation for managing your dbt™ jobs effectively in production environments.

Next, we'll explore different [schedule types and triggers](/app-help/guides/paradime-101/running-dbt-in-production-with-bolt/understanding-schedule-types-and-triggers) to give you more flexibility in managing your dbt™ workflows.


# Understanding schedule types and triggers

Scheduling is a crucial part of managing your data workflows. Paradime offers flexible options to ensure your jobs run exactly when you need them and in the most efficient manner.

**Estimated completion time:** 10 minutes

{% hint style="warning" %}

#### Prerequisites

* Access to the [Bolt application](https://www.paradime.io/pricing) in Paradime
* A [Production Connection](/app-help/guides/paradime-101/getting-started-with-your-paradime-workspace/setting-up-data-warehouse-connections) to your data warehouse (AKA Schedule Connection)
* [At least one dbt™ model materialized in your dbt™ project](https://docs.paradime.io/app-help/guides/paradime-101/running-dbt-in-production-with-bolt/pages/xnzMpxvtdDSGzR3giyID#id-4.-materialize-your-dbt-tm-model)
  {% endhint %}

### What You'll Learn

In this guide, you'll learn about:

1. [Different types of Bolt schedules ](#id-1.-schedule-types)(Standard, Deferred, and Turbo CI) and when to use them
2. Various [schedule triggers](#id-2.-schedule-triggers) (Scheduled run, On Run Completion, On Merge, Bolt API) and when to use them)
3. [Schedule statuses](#id-3.-schedule-statuses) and their meanings

***

### 1. Schedule Types

Paradime supports three main schedule types, each designed for specific use cases:

<table><thead><tr><th width="150">Type</th><th>Description</th><th>Best for</th></tr></thead><tbody><tr><td><strong>Standard</strong></td><td>Runs your dbt™ job directly in the production environment</td><td>Regular production runs, daily or hourly updates</td></tr><tr><td><strong>Deferred</strong></td><td>Runs your job in a separate environment before merging changes to production</td><td>Testing changes before deploying to production, CI/CD workflows</td></tr><tr><td><strong>Turbo</strong> CI</td><td>Performs rapid, incremental runs of only changed models</td><td>Fast feedback on model changes, efficient CI processes</td></tr></tbody></table>

For more details on setting up each type, refer to our [advanced scheduling documentation](/app-help/documentation/bolt/managing-schedules/schedule-configurations).

***

### 2. Schedule Triggers

{% embed url="<https://youtu.be/YCvU_8qzjWg>" %}

Paradime supports four main schedule triggers:

<table><thead><tr><th width="196">Schedule Trigger</th><th>What it does</th><th width="217">How it works</th><th>Best for</th></tr></thead><tbody><tr><td><strong>Scheduled Run</strong></td><td>Runs your job at specific times and frequencies</td><td>Uses cron syntax to set the schedule</td><td>Regular, time-based job execution (e.g., daily reports, weekly updates)</td></tr><tr><td><strong>On Run Completion</strong></td><td>Triggers your job after another specified job finishes</td><td>You select a "parent" job, and this job will start once the parent completes</td><td>Creating dependencies between jobs (e.g., running a summary job after individual data updates)</td></tr><tr><td><strong>On Merge</strong></td><td>Runs your job when a pull request is merged into a specified branch</td><td>Connects to your Git repository and triggers based on merge events.<br><br>See <a href="/pages/v0CmoMvcol2UIWv3VnSQ#defer-to-production-configuration">Defer to Production</a> documentation for details.</td><td>Continuous Deployment workflows, ensuring your data pipeline updates with your code changes</td></tr><tr><td><strong>Bolt API</strong></td><td>Allows you to trigger jobs programmatically from your existing data pipelines</td><td>Provides API endpoints to start, stop, or check job status<br><br>See <a href="https://docs.paradime.io/app-help/documentation/bolt/managing-schedules/bolt-api">Bolt API documentation</a> for details.<br></td><td>Integrating Paradime jobs with external systems or custom workflows</td></tr></tbody></table>

***

### 3. Schedule Statuses

As you manage multiple Bolt schedules, you'll encounter the following statuses:

| Status      | Description                                                                         |
| ----------- | ----------------------------------------------------------------------------------- |
| ✅ Success   | The schedule has completed successfully                                             |
| 🚫 Canceled | The schedule was manually canceled or stopped before completion                     |
| ❌ Error     | The schedule encountered an error during execution (AKA the scheduled run "failed") |
| ⏸️ Paused   | The schedule has been temporarily paused                                            |
| 🕐 No runs  | The schedule has yet to execute and/or it is not configured to execute.             |

***

{% hint style="info" %}

#### Related Documentation

* [How to use Trigger Type "Defer To Production"](/app-help/documentation/bolt/managing-schedules/schedule-configurations#defer-to-production-configuration)
* [Using --defer in Paradime](/app-help/concepts/paradime-fundamentals/dbt-tm-defer-to-production)
* [Bolt API and Webhooks](/app-help/documentation/bolt/bolt-api)
* [TurboCI](/app-help/documentation/bolt/ci-cd/turbo-ci)
  {% endhint %}

***

### Summary

You've learned about the different types of Bolt schedules, various trigger methods, and schedule statuses. This knowledge will help you choose the right schedule type and trigger for your specific use cases, and understand the status of your running schedules.

Next, we'll explore how to [view run history and analytics for your Bolt schedules](/app-help/guides/paradime-101/running-dbt-in-production-with-bolt/viewing-run-history-and-analytics), which will help you monitor and optimize your dbt™ workflows.


# Viewing Run History and Analytics

Bolt provides comprehensive tools for monitoring and analyzing your dbt™ schedules. This guide will walk you through viewing run logs, understanding the Bolt Schedules list, and exploring detailed analytics for your schedules.

**Estimated completion time:** 15 minutes

{% hint style="warning" %}

#### Prerequisites

* Access to the [Bolt application](https://www.paradime.io/pricing) in Paradime
* At least one configured [Bolt schedule](/app-help/guides/paradime-101/running-dbt-in-production-with-bolt/creating-bolt-schedules)
  {% endhint %}

### What you'll learn

In this guide, you'll learn how to:

1. [Navigate the Bolt Schedules overview](#id-1.-bolt-schedules-overview)
2. [Access and interpret Bolt Schedule detail views](#id-2.-bolt-schedule-detail-views)
3. [Understand run history, logs, and artifacts](#id-2.1-run-history)

***

### 1. Bolt Schedules Overview

{% embed url="<https://youtu.be/gYOI1hqljzU>" %}

The Bolt home screen displays a summarized analysis of all your configured schedules, both UI-based and [YAML-based](/app-help/documentation/bolt/creating-schedules#yaml-based-schedules). Here's what you can see at a glance:

<table><thead><tr><th>Field</th><th width="319">Description</th><th>Example</th></tr></thead><tbody><tr><td><strong>Name</strong></td><td>Name of the schedule</td><td><code>hourly scheduled run</code></td></tr><tr><td><strong>Status</strong></td><td>Current status of the schedule</td><td><code>Success</code></td></tr><tr><td><strong>Owner</strong></td><td>The schedule's owner (if configured)</td><td><code>john@acme.com</code></td></tr><tr><td><strong>Cron Description</strong></td><td>Human-readable description of the schedule's run time (UTC)</td><td>At 00:00, every day (UTC)</td></tr><tr><td><strong>Cron Configuration</strong></td><td>The cron syntax for the scheduled run</td><td>@daily</td></tr><tr><td><strong>Last Run</strong></td><td>Date and time of the most recent execution</td><td>August 22, 2024 at 5:00 PM PDT</td></tr><tr><td><strong>Next Run</strong></td><td>Anticipated date and time for the next execution</td><td>August 23, 2024 at 5:00 PM PDT</td></tr><tr><td><strong>Until Next Run</strong></td><td>Time remaining until the next scheduled runT</td><td>20 hours, 38 minutes, 52 seconds</td></tr><tr><td><strong>Trigger Type</strong></td><td>How the schedule is triggered (On schedule, On Run Completion)</td><td><code>standard</code></td></tr><tr><td><strong>On Completion Configuration</strong></td><td>Details of On Run Completion configuration (if applicable)</td><td><code>schedule base run in workspace irishtrooper on status passed, failed</code></td></tr></tbody></table>

### 2. Bolt Schedule Detail Views

The Bolt Schedule Detail View provides in-depth insights into individual dbt™ schedules, including configuration, run history, logs, and artifacts.

How to to access:

1. **Schedules Overview:** Navigate to the main page displaying all your configured schedules.
2. **Schedule Selection:** Click on a specific schedule name to view its detailed information and run history.

#### 2.1: Run History

{% embed url="<https://youtu.be/vjvDMwtTfWE>" %}

In the Run History section, you can view all executions of a specific scheduled run:

| Field                 | Description                                                   | Example                     |
| --------------------- | ------------------------------------------------------------- | --------------------------- |
| **Status**            | Status of the specific runID                                  | `Success`                   |
| **Trigger**           | If the runID was manually triggered manually or automatically | `Manual from john@acme.com` |
| **Branch and commit** | The branch name and commit SHA used when running the schedule | `main #ce34f`               |
| **Last Run**          | Date and time of when the runID was executed                  | `last month`                |
| **Duration**          | How long the run took to complete                             | `25 seconds`                |
| **RunID**             | Unique identifier of the run                                  | `13403`                     |

#### 2.2: Logs and Artifacts

{% embed url="<https://youtu.be/ynkNmpBIRW8>" %}

Every time a dbt™ command executes in a Bolt schedule, dbt Core™ generates a set of artifacts like manifest.json, catalog.json, run\_results.json, and sources.json. These tools help analyze, troubleshoot, and optimize your Bolt schedules.

**How to access:**

* Click on a specific run from a Bolt schedule's Run History
* Scroll down to Logs and Artifacts section

Let's explore the key components:

#### 2.3: Run Logs

By clicking on a specific command executed in your bolt schedule (e.g., `dbt run`) you'll have access to various run logs:

| Run Log Type     | Description                                                                                                                                                                                                               | Use Case                                                            |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------- |
| **Summary Logs** | [DinoAI-generated](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/dinoai-accelerating-your-analytics-engineering-workflow) overview of dbt command execution, including warnings and potential fixes | Quick assessment of run health and identification of common issues  |
| **Console Logs** | Detailed, chronological record of all operations                                                                                                                                                                          | Detailed troubleshooting and understanding of the execution process |
| **Debug Logs**   | Extensive details including system-level operations and dbt™ internals                                                                                                                                                    | In-depth problem solving and performance tuning for complex issues  |

#### 2.4: Source Freshness

When your Bolt Schedule contains the command 'dbt source freshness', Paradime will provide the state of each of your sources. This helps you determine the status of your source data and whether it aligns with your SLAs.

<figure><img src="/files/nmtYUR1or3dZsH0imCgO" alt=""><figcaption></figcaption></figure>

#### 2.5: Artifacts

During a Bolt schedule execution, dbt™ generates various files including run SQL files, compiled SQL files, manifest files, and JSON files. These artifacts provide insights into what was executed and how, enabling you to analyze the output of your dbt runs in detail.

<figure><img src="/files/CO7XHjEJZFlqx7hSxdS6" alt=""><figcaption></figcaption></figure>

***

{% hint style="info" %}

#### Related Documentation

* [View Bolt Run Logs](/app-help/documentation/bolt/managing-schedules/viewing-run-log-history)
* [Viewing Bolt Schedule Source Freshness](/app-help/documentation/bolt/managing-schedules/analyzing-run-details/configuring-source-freshness)
  {% endhint %}

***

### Summary

You've learned how to navigate the Bolt Schedules overview, access detailed information about specific schedules, and interpret run history, logs, and artifacts. This knowledge will help you effectively monitor and analyze your dbt™ schedules, enabling you to optimize your data workflows and quickly troubleshoot any issues.

Next, we'll explore [how to set up notifications for your Bolt schedules](/app-help/guides/paradime-101/running-dbt-in-production-with-bolt/setting-up-notifications), ensuring you stay informed about the status of your data pipelines.


# Setting Up Notifications

Staying informed about your Bolt Schedule statuses is crucial for maintaining a smooth data pipeline. Bolt allows you to send granular, schedule-level notifications via Slack or email to alert users in your organization when a scheduled run passes or fails.

**Estimated completion time:** 10 minutes

{% hint style="info" %}

#### Prerequisites

* Access to the [Bolt application](https://www.paradime.io/pricing) in Paradime
* At least one configured [Bolt schedule](/app-help/guides/paradime-101/running-dbt-in-production-with-bolt/creating-bolt-schedules)
* [Paradime Slack integration set up](/app-help/documentation/integrations/notifications/slack) (if planning to use Slack notifications)
  {% endhint %}

### What you'll learn

In this guide, you'll learn how to:

1. [Configure Bolt Schedule notifications](#id-1.-configuring-bolt-schedule-notifications)
2. [Set up Slack and Email notifications](#id-2.-using-schedule-notifications)

***

### 1. Configuring Bolt Schedule Notifications

{% embed url="<https://youtu.be/SsT0SVyvv-4>" %}

To set up notifications for your Bolt Schedule:

1. Navigate to the Bolt home screen
2. Select a specific Bolt Schedule
3. Click `Edit` within your [Bolt schedule's detail view](/app-help/guides/paradime-101/running-dbt-in-production-with-bolt/viewing-run-history-and-analytics#bolt-schedule-detail-views)
4. Scroll down to the `notification settings` and configure your schedule as needed:

| Notification Setting         | Description                                                  | Example                                                                  |
| ---------------------------- | ------------------------------------------------------------ | ------------------------------------------------------------------------ |
| **Slack notify on**          | Specifies when Slack notifications are sent                  | Send Slack notification when a Bolt schedule `passed` and/or `failed`    |
| **Slack channels and users** | List of Slack channels and/or users to receive notifications | <p><code>@john</code><br><code>#scheduledruns</code></p>                 |
| **Email notify on**          | Specifies when email notifications are sent                  | Send an email notification when a Bolt schedule `passed` and/or `failed` |
| **Email notify**             | Email addresses to receive notifications                     | `john@acme.com`                                                          |

5. Click `Publish`to save your changes

{% hint style="info" %}
If you're Bolt Scheduled already has a [trigger type](/app-help/guides/paradime-101/running-dbt-in-production-with-bolt/understanding-schedule-types-and-triggers#schedule-triggers) configured, your notifications will send out following the next run. However, you can manually execute a bolt run from your schedule's detail view to verify your notification(s) setup immediately.
{% endhint %}

<figure><img src="/files/ezNdnAMSrmkwlMTv5uZW" alt=""><figcaption></figcaption></figure>

### 2. Using Schedule notifications

Following an execution of your Bolt schedule, notifications will be sent to the specified platforms and people. By clicking "Open in Paradime" in the notification, you'll be redirected to your Bolt Schedule's view for further investigation.

#### Slack Notification

<figure><img src="/files/lp0awuOxvSjjpfwVrg5f" alt="" width="563"><figcaption></figcaption></figure>

#### Email Notification

<figure><img src="/files/8NwK1TE37avCBRnd2p5E" alt="" width="563"><figcaption></figcaption></figure>

***

{% hint style="info" %}

#### Related Documentation

* [Bolt Schedule Notifications](/app-help/documentation/bolt/creating-schedules/notification-settings)
* [Paradime Slack Integration Setup](/app-help/documentation/integrations/notifications/slack)
  {% endhint %}

***

### Summary

You've learned how to set up and configure notifications for your Bolt schedules, including both Slack and email notifications. This will help you stay informed about the status of your data pipelines and quickly respond to any issues that arise.

With this knowledge, you're now equipped to effectively manage, monitor, and maintain your dbt™ schedules using Bolt in Paradime.


# Debugging Failed Runs

Learn how to efficiently debug failed dbt™ runs in Paradime using Bolt.

When running dbt™ in production, it's crucial to quickly identify and resolve failed runs. This guide will walk you through the process of debugging failed dbt™ runs using Bolt's comprehensive logging and debugging tools.

**Estimated completion time:** 15 minutes

{% hint style="warning" %}
**Prerequisites**

* Access to the [Bolt application](https://www.paradime.io/pricing) in Paradime
* At least one configured Bolt schedule
* Basic understanding of dbt™ commands and SQL
  {% endhint %}

{% hint style="info" %}
**What you'll learn**

In this guide, you'll learn how to:

1. [Identify failed runs](#id-1.-identifying-failed-runs)
2. [Access and interpret run logs](#id-2.-access-and-interpret-run-logs)
3. [Debug and resolve common issues](#id-3.-debug-and-resolve-common-issues)
   {% endhint %}

{% embed url="<https://youtu.be/yJ1z2m-fRAw>" %}

***

### 1. Identifying Failed Runs

There are two main ways to identify failed runs:

**Method 1: Notifications**

Set up [Slack or email notifications](/app-help/guides/paradime-101/running-dbt-in-production-with-bolt/setting-up-notifications) to receive immediate alerts when a run fails. This is the recommended approach for production environments.

{% hint style="info" %}
See [Setting Up Notifications](/app-help/guides/paradime-101/running-dbt-in-production-with-bolt/setting-up-notifications) for details on configuring alerts.
{% endhint %}

**Method 2: Bolt UI**

Navigate to the [Bolt home screen](https://app.paradime.io/bolt/) and check the "Status" column in the Bolt Schedule List. Failed runs are marked with a "Error" status indicator.

{% hint style="info" %}
See [Viewing Run History and Analytics](/app-help/guides/paradime-101/running-dbt-in-production-with-bolt/viewing-run-history-and-analytics) for more details on using the Bolt UI.
{% endhint %}

***

### 2. Access and interpret run logs

Once you've identified a failed run, access the logs through the Bolt Schedule Detail Views:

1. Click on the failed Bolt Schedule (one marked with a "Error" status indicator)
2. Navigate to the **Run History** section
3. Select the failed run (one marked with a "Error" status indicator)
4. Scroll to the **Logs and Artifacts** section
5. Click on the executed command that failed (ex. `dbt run`)

Bolt Provides three types of logs:

| Run Log Type     | Description                                                                                                                                                          | Use Case                                 |
| ---------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------- |
| **Summary Logs** | [DinoAI-generated](/app-help/guides/paradime-101/getting-started-with-the-paradime-ide/dinoai-accelerating-your-analytics-engineering-workflow) overview of failures | Quick assessment of issues               |
| **Console Logs** | Detailed execution record                                                                                                                                            | Finding specific errors and compiled SQL |
| **Debug Logs**   | System-level operations                                                                                                                                              | Deep technical troubleshooting           |

{% hint style="info" %}
**Console logs** are typically the most useful for debugging as they show errors, warnings, and the compiled SQL code that failed in production.
{% endhint %}

***

### 3. Debug and resolve common issues

Follow these steps to debug a failed run:

1. **Review Summary Logs**
   * Check the AI-generated overview
   * Note any suggested fixes

{% code title="Summary Logs Example" %}

```
Command executed 6 models:
- 5 models passed
- 1 model failed
- Error in "fct_fantasy_point_leaders" model
- Invalid identifier 'TOTAL_FANTASY_POINTS_PP' on line 9
```

{% endcode %}

2. **Review Console Logs**
   * Locate error messages and warnings using the "jump to" feature
   * Click on the link for compiled SQL code
   * Review the execution flow and timing

<figure><img src="/files/aaX67SAelWr82h75Jegl" alt=""><figcaption></figcaption></figure>

3. **Test and Fix**
   * Copy the compiled SQL code from console logs
     * Test the SQL:
       * Directly in your data warehouse, or
       * In the Code IDE [scratchpad](/app-help/documentation/code-ide/additional-features/scratchpad)
     * Fix common issues:
       * Invalid column names
       * Missing model references
       * SQL syntax errors
       * Data type mismatches

<figure><img src="/files/F5Sq5TAP26RRYyFOtXNN" alt=""><figcaption></figcaption></figure>

After testing the compiled SQL code against your data warehouse, you'll be better equipped to resolve common issues with failed scheduled runs.

***

{% hint style="info" %}
**Related Documentation**

* [Setting Up Notifications](/app-help/documentation/bolt/creating-schedules/notification-settings)
* [Viewing Run History and Analytics](/app-help/documentation/bolt/managing-schedules/viewing-run-log-history)
  {% endhint %}

***

#### Summary

You've learned how to identify failed dbt™ runs, access and interpret different types of logs, and systematically debug and resolve issues. This knowledge will help you maintain reliable data pipelines and quickly resolve any failures that occur.

Next, consider exploring the related documentation to handle more complex failure scenarios.


# Getting started with the DinoAI background agent

DinoAI Agent runs autonomously in the background — executing dbt™ models, querying your warehouse, and opening PRs without you having to be in the IDE. Before your first run, you need to connect three things.

{% hint style="warning" icon="plug-circle-plus" %}
**New to Paradime? Start by setting up your workspace before using DinoAI.**

> ⚠️ **Note:** A CodeIDE connection is required during workspace onboarding — even if you plan to use the agent only. This requirement will be removed in a future update.

Follow our step-by-step guide to get started: [Setting up your Paradime workspace →](https://docs.paradime.io/app-help/guides/paradime-101/getting-started-with-your-paradime-workspace/creating-a-workspace)
{% endhint %}

{% hint style="info" %}
**Before you start, make sure you have the following:**

* Admin access to your Paradime workspace
* GitHub organisation owner access (or someone on your team who does)
* Data warehouse credentials for DinoAI — we recommend a dedicated service account with scoped permissions

Watch the setup tour below, or follow the step-by-step docs.
{% endhint %}

{% @arcade/embed url="<https://app.arcade.software/share/2zqa5J7w53ciLDdxBgAl>" flowId="2zqa5J7w53ciLDdxBgAl" %}

### 1. Connect Slack

{% hint style="warning" %}
If Slack is already connected, disconnect and reconnect to pick up the new background agent permissions.
{% endhint %}

DinoAI communicates through Slack — progress updates, results, and errors all land in your connected workspace.

*Workspace Settings > Integrations > Paradime Slack App > Connect*

You'll be redirected to Slack to authorise the app. Sign in and approve the required permissions. Learn more [here](/app-help/documentation/integrations/notifications/slack).

***

### 2. Install the GitHub app

DinoAI needs GitHub access to push branches and open PRs against your connected repo. Learn more [here](/app-help/documentation/integrations/ci-cd/github).

*Workspace Settings > Integrations > GitHub > Connect*

1. Authenticate with GitHub when redirected.
2. Select the repositories to connect — include the repo tied to your Paradime workspace.
3. Click **Install and authorize**.

{% hint style="warning" %}
You must be a **GitHub organisation owner** to install the app. If you're not, click **Authorize & Request** — this sends an automated approval request to your org owner.
{% endhint %}

***

### 3. Set up the agent environment

This is the warehouse connection DinoAI uses at runtime when executing models or running queries on your behalf.

*Account Settings > Connections > DinoAI Background Agent Environment*

Select your warehouse and follow the configuration guide for your platform:

{% hint style="info" %}
**We recommend creating a dedicated service user with its own permissions for the background agent, with read-only access to your production database.**
{% endhint %}

***

{% hint style="success" %}
**You are ready!**

Once all three are connected, DinoAI can run fully in the background. Head to Agent Mode to kick off your first task.
{% endhint %}

{% hint style="info" icon="t-rex" %}
**Customise how DinoAI operates with `.dinorules`**

Want DinoAI to follow your team's conventions, coding standards, or workflow preferences? Add a `.dinorules` file to the root of your repository and merge it into your default branch — DinoAI will pick it up automatically at the start of every session.

\
[**Learn how to configure `.dinorules` →**](https://docs.paradime.io/app-help/documentation/dino-ai/dino-rules#example-.dinorules-file-configuration)
{% endhint %}

### Start your first DinoAI session

Setup done. **Time to wake DinoAI up 🦖**. Go to Slack and add the Paradime bot to the channel where you want DinoAI to operate.

{% hint style="info" %}
To add an DinoAI to a Slack channel, either use the command `/invite @paradime` within that channel or click the channel name, go to the **Integrations** tab, select **Add an App**, and add Paradime.
{% endhint %}

{% hint style="warning" %}
Always use `@paradime` when sending messages to the DinoAI agent in Slack. Without it, the agent won't see or respond to your message.
{% endhint %}

{% @arcade/embed url="<https://app.arcade.software/share/oM4ejYcBMzW7joQtPSSC>" flowId="oM4ejYcBMzW7joQtPSSC" %}

Then send your first message to check it's listening:

> @paradime Hey Dino, are you there?

If everything is connected, DinoAI will respond and confirm it's ready to take on tasks.

From there, give it a real first task — something like:

> *"Dino, document all models in the marts folder that are missing descriptions."*

> *"Dino, create a new model that joins orders and customers on customer\_id."*

> *"Dino, run a quick check on our staging models and flag anything that looks off."*

DinoAI will pick it up, work in the background, and report back in the channel when it's done.


# Snowflake DinoAI Agent Key-Pair Setup

Set up Snowflake key-pair authentication for the DinoAI Background Agent by generating RSA keys, creating a service user, granting access, and configuring the connection in Paradime.

## Snowflake DinoAI Agent Key-Pair Setup

{% hint style="info" %}
We recommend creating a dedicated Snowflake service user for the DinoAI agent environment, with the minimum permissions required to read and write to your database.
{% endhint %}

{% hint style="info" %}
**IP Restrictions** If your Snowflake account uses network policies, make sure to allowlist the Paradime IPs for your data region.\
\
👉 See: [Paradime IP addresses](https://docs.paradime.io/app-help/developers/ip-restrictions)
{% endhint %}

***

#### Step 1. Generate the Key Pair

Run the following commands in your terminal to generate an RSA private key and extract the public key.

**Option A — Encrypted private key (recommended)**

```bash
# Generate a 2048-bit RSA private key encrypted with AES-256
openssl genrsa 2048 | openssl pkcs8 -topk8 -v2 aes-256-cbc -inform PEM -out dinoai_rsa_key.p8

# Extract the public key
openssl rsa -in dinoai_rsa_key.p8 -pubout -out dinoai_rsa_key.pub
```

You'll be prompted to set a passphrase. Keep it — you'll need it when configuring the connection in Paradime.

**Option B — Unencrypted private key**

```bash
# Generate an unencrypted 2048-bit RSA private key
openssl genrsa 2048 | openssl pkcs8 -topk8 -nocrypt -inform PEM -out dinoai_rsa_key.p8

# Extract the public key
openssl rsa -in dinoai_rsa_key.p8 -pubout -out dinoai_rsa_key.pub
```

**Get the public key value**

```bash
cat dinoai_rsa_key.pub
```

{% hint style="info" %}
Copy only the content **between** the header and footer lines — not the lines themselves.\
i.e. everything between `-----BEGIN PUBLIC KEY-----` and `-----END PUBLIC KEY-----`
{% endhint %}

***

#### Step 2. Create the Snowflake Service User

Run the SQL below in a Snowflake worksheet to create a dedicated service user, role, and warehouse.

**Create user, role, and warehouse**

```sql
use role securityadmin;

-- Create warehouse (skip if one already exists)
create warehouse if not exists transforming
    warehouse_size = xsmall
    auto_suspend = 60
    auto_resume = true
    initially_suspended = true;

-- Create the transformer role and grant warehouse access
create role if not exists transformer;
grant all on warehouse transforming to role transformer;

-- Create the DinoAI service user
create user paradime_dinoai_user
    password = '<generate_a_strong_password>'
    default_warehouse = transforming
    default_role = transformer;

-- Assign the role to the user
grant role transformer to user paradime_dinoai_user;
```

**Grant database permissions**

The DinoAI agent needs three levels of access across your Snowflake databases:

| Database                                   | Access level     | Why                                                                         |
| ------------------------------------------ | ---------------- | --------------------------------------------------------------------------- |
| **Dev database** (e.g. `dev`)              | **Read + Write** | DinoAI creates and modifies tables/views during background agent sessions   |
| **Production database** (e.g. `analytics`) | **Read only**    | DinoAI reads prod models to understand your data and generate accurate code |
| **Source database** (e.g. `raw`)           | **Read only**    | DinoAI inspects raw source tables to understand upstream data structures    |

Replace the database names below with your actual Snowflake databases.

```sql
use role sysadmin;

-- -------------------------
-- DEV DATABASE — read + write
-- DinoAI builds and modifies objects here during agent sessions
-- -------------------------
grant all on database <your_dev_database> to role transformer;
grant all on all schemas in database <your_dev_database> to role transformer;
grant all on future schemas in database <your_dev_database> to role transformer;
grant all on all tables in database <your_dev_database> to role transformer;
grant all on future tables in database <your_dev_database> to role transformer;
grant all on all views in database <your_dev_database> to role transformer;
grant all on future views in database <your_dev_database> to role transformer;

-- -------------------------
-- PRODUCTION DATABASE — read only
-- DinoAI reads prod models to understand your existing data
-- -------------------------
grant usage on database <your_prod_database> to role transformer;
grant usage on all schemas in database <your_prod_database> to role transformer;
grant usage on future schemas in database <your_prod_database> to role transformer;
grant select on all tables in database <your_prod_database> to role transformer;
grant select on future tables in database <your_prod_database> to role transformer;
grant select on all views in database <your_prod_database> to role transformer;
grant select on future views in database <your_prod_database> to role transformer;

-- -------------------------
-- SOURCE DATABASE — read only
-- DinoAI inspects raw source tables to understand upstream data
-- -------------------------
grant usage on database <your_source_database> to role transformer;
grant usage on all schemas in database <your_source_database> to role transformer;
grant usage on future schemas in database <your_source_database> to role transformer;
grant select on all tables in database <your_source_database> to role transformer;
grant select on future tables in database <your_source_database> to role transformer;
grant select on all views in database <your_source_database> to role transformer;
grant select on future views in database <your_source_database> to role transformer;
```

{% hint style="info" %}
If your source and production data live in the same database but different schemas, replace the database-level grants with schema-level grants.\
e.g. `grant select on all tables in schema <your_database>.<your_schema> to role transformer;`
{% endhint %}

***

#### Step 3. Assign the Public Key to the User

Once the user is created, assign the RSA public key to enable key-pair authentication.

```sql
use role securityadmin;

-- Paste the public key content here (no header/footer lines, key must be on a single line)
alter user paradime_dinoai_user set rsa_public_key='MIIBIjANBgkqhkiG9w0BAQEF...';
```

**Verify the key was assigned**

```sql
desc user paradime_dinoai_user;
```

Check that the `RSA_PUBLIC_KEY` column is populated.

**(Optional) Verify the fingerprint**

Cross-check that the key in Snowflake matches your local key.

```bash
# Get the SHA-256 fingerprint of your local public key
openssl rsa -pubin -in dinoai_rsa_key.pub -outform DER | openssl dgst -sha256 -binary | openssl enc -base64
```

```sql
-- Compare against what Snowflake stored
select name, rsa_public_key_fp
from table(information_schema.users())
where name = 'PARADIME_DINOAI_USER';
```

***

#### Step 4. Configure the Connection in Paradime

1. Click **Settings** in the top menu bar
2. Click **Connections** in the left sidebar
3. Click **Add New** next to **DinoAI Background Agent Environment**
4. Select **Snowflake** as the connection type
5. Choose **Key-Pair Authentication**

Fill in the fields as follows:

| Field                   | Description                                       | Example                                 |
| ----------------------- | ------------------------------------------------- | --------------------------------------- |
| Profile                 | Profile name from your `dbt_project.yaml`         | `dbt-snowflake`                         |
| Target                  | Target name for the DinoAI connection             | `dinoai`                                |
| Account                 | Your Snowflake account identifier                 | `vj71689.eu-west-2.aws`                 |
| Role                    | Role assigned to the service user                 | `transformer`                           |
| Database                | The **dev** database DinoAI will build objects in | `dev`                                   |
| Warehouse               | Virtual warehouse for agent sessions              | `transforming`                          |
| Username                | Service user created in Step 2                    | `paradime_dinoai_user`                  |
| Private Key             | Full private key including header/footer lines    | `-----BEGIN ENCRYPTED PRIVATE KEY-----` |
| Passphrase *(optional)* | Passphrase set when generating the key            | `passphrase_xyz`                        |
| Schema                  | Default schema for dbt objects at runtime         | `dbt_dinoai`                            |
| Threads                 | Number of concurrent threads                      | `8`                                     |

{% hint style="warning" %} When pasting the **Private Key**, you must include the full header and footer lines:

```
-----BEGIN ENCRYPTED PRIVATE KEY-----
< key content >
-----END ENCRYPTED PRIVATE KEY-----
```

Omitting them will cause the connection to fail. {% endhint %}

***

#### Step 5. Allowlist Paradime IPs (if using network policies)

If your Snowflake account has IP restrictions, allowlist the Paradime IPs for your data region.

```sql
-- Add Paradime IPs to an existing network policy
alter network policy <your_policy_name>
    set allowed_ip_list = ('<paradime_ip_1>', '<paradime_ip_2>');

-- Apply the policy to the DinoAI service user
alter user paradime_dinoai_user set network_policy = <your_policy_name>;
```

👉 Full IP list [here](/app-help/developers/ip-restrictions).

***

#### Troubleshooting

**JWT token errors on connection test**

* Run `DESC USER paradime_dinoai_user;` and verify `RSA_PUBLIC_KEY` is populated
* Make sure the private key in Paradime includes the header/footer lines
* If using an encrypted key, confirm the passphrase is correct

**Permission denied during agent runs**

* Verify grants were run as `SYSADMIN`, not `SECURITYADMIN`
* Check future grants are in place — new schemas/tables won't inherit permissions without them
* Run `SHOW GRANTS TO ROLE transformer;` to audit what the role can access
* If DinoAI can't read source or prod data, double-check read grants on those databases

**Connection timeout / network errors**

* Ensure Paradime IPs are allowlisted in your Snowflake network policy (Step 5)
* Verify the account identifier format matches [Snowflake's documentation](https://docs.snowflake.com/en/user-guide/admin-account-identifier)


# Migrating Guide - dbt™ cloud to Paradime

dbt Cloud™ Importer in Paradime: Import dbt Cloud™ projects into Paradime with one click to simplify project management and integration.

This is comprehensive guide to migrate from dbt Cloud™️ to Paradime. This guide covers the following important steps:

* Understand the translation of key terms between both platforms
* Guide to transferring dbt™️ jobs to Paradime - it takes only 3-clicks 🔥


# Translation of Key Terms

dbt Cloud™️ to Paradime migration guide with the translation of key terms between both the platforms.

This guide maps terminology and concepts between dbt Cloud<sup>TM</sup> and [Paradime.io](http://paradime.io) to help teams migrate or understand the differences between platforms.

***

### Core Platform Components

| dbt Cloud™️                    | paradime.io           | Notes                                                                                                                                      |
| ------------------------------ | --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------ |
| dbt Cloud<sup>TM</sup>         | Paradime              | The overall platform name                                                                                                                  |
| dbt Platform<sup>TM</sup>      | Paradime Platform     | The complete analytics operating system                                                                                                    |
| dbt Cloud<sup>TM</sup> Project | Paradime Workspace    | The namespace containing the data warehouse connections, integrations, users and dbt code to enable development and execution of pipelines |
| Studio IDE                     | Code IDE              | Browser-based development environment                                                                                                      |
| Orchestration                  | Bolt Scheduler / Bolt | Overall scheduling system                                                                                                                  |
| dbt<sup>TM</sup> Catalog       | Paradime Catalog      | The interface to explore metadata, documentation, and lineage for your data (+ integrations with BI Tool - Paradime only)                  |

### Orchestration & Scheduling

| dbt Cloud™️   | paradime.io             | Notes                                                  |
| ------------- | ----------------------- | ------------------------------------------------------ |
| Job           | Bolt Schedule           | Scheduled execution of dbt<sup>TM</sup> commands       |
| Orchestration | Bolt Scheduler / Bolt   | Overall scheduling system                              |
| Run           | Schedule Run / Bolt Run | Single execution of a job/schedule                     |
| Deploy        | Deploy                  | Deployment mechanism for executing commands            |
| Run History   | Run History             | Historical execution data                              |
| N/A           | paradime\_schedules.yml | YAML file for defining schedules as code (git-tracked) |

***

### Schedule/Job Types

| dbt Cloud™️       | paradime.io       | Notes                                              |
| ----------------- | ----------------- | -------------------------------------------------- |
| Deployment Job    | Standard Schedule | Regular scheduled runs on set intervals            |
| CI Job            | Turbo CI Schedule | Runs on pull request events                        |
| N/A               | Deferred Schedule | Optimized runs using manifest / sources comparison |
| Scheduled Trigger | Scheduled Run     | Time-based execution (cron)                        |
| API Trigger       | Bolt API Trigger  | API-based triggering                               |
| On Merge          | On Merge Trigger  | Runs when PR is merged (CD)                        |
| Job Completion    | On Run Completion | Runs after another execution completes             |

***

### Environments

| dbt Cloud™️                         | paradime.io            | Notes                                   |
| ----------------------------------- | ---------------------- | --------------------------------------- |
| Development Environment             | Code IDE Environment   | IDE workspace for individual developers |
| Deployment / Production Environment | Production Environment | Where jobs execute                      |
| Staging Environment                 | -                      | Pre-production validation layer         |

***

### Development & IDE Features

| dbt Cloud™️             | paradime.io             | Notes                                                             |
| ----------------------- | ----------------------- | ----------------------------------------------------------------- |
| Studio IDE              | Code IDE                | Cloud-based development interface                                 |
| Command Bar             | Integrated Terminal     | Execute dbt<sup>TM</sup> commands                                 |
| N/A                     | DinoAI                  | AI-powered coding assistant                                       |
| N/A                     | Scratchpad              | Temporary SQL file for exploration                                |
| Git Integration         | Git Lite / Git Advanced | Version control integration                                       |
| Defer to Production     | Defer to Prod           | Development using production artifacts                            |
| Development Credentials | Code IDE Credentials    | Individual developer warehouse access                             |
| N/A                     | Catalog                 | Preview documentation in IDE                                      |
| Lineage                 | Lineage Preview         | Preview lineage during development                                |
| Preview                 | Data Explorer           | Preview dbt<sup>TM</sup> model and SQL results during development |

***

### Documentation & Discovery

| dbt Cloud™️           | paradime.io           | Notes                                               |
| --------------------- | --------------------- | --------------------------------------------------- |
| dbt Docs<sup>TM</sup> | dbt Docs<sup>TM</sup> | Generated documentation for dbt<sup>TM</sup> assets |
| dbt Explorer          | Paradime Lineage      | Interactive lineage graph                           |
| Catalog               | Data Catalog          | Metadata and documentation repository               |
| N/A                   | Looker Assets         | Looker-specific lineage tracking                    |
| N/A                   | Tableau Assets        | Tableau-specific lineage tracking                   |
| N/A                   | Power BI Assets       | Power BI integration                                |
| N/A                   | ThoughtSpot Assets    | ThoughtSpot integration                             |

### CI/CD & State Management

| dbt Cloud™️            | paradime.io                        | Notes                                |
| ---------------------- | ---------------------------------- | ------------------------------------ |
| Slim CI                | Turbo CI                           | Build only modified models           |
| Compare against        | Deferred Schedule                  | State comparison for slim builds     |
| Production State       | Production Manifest                | Source of truth for comparison       |
| Continuous Integration | Turbo CI                           | Automated testing on PRs             |
| Continuous Deployment  | On Merge / Custom CD               | Automated deployment                 |
| N/A                    | Column Level Lineage Diff Analysis | Visual comparison of lineage changes |

***

### Git & Version Control

| dbt Cloud™️     | paradime.io  | Notes                     |
| --------------- | ------------ | ------------------------- |
| Version Control | Git Lite     | Simplified git workflow   |
| N/A             | Git Advanced | Full git capabilities     |
| N/A             | GitLens      | Git history visualization |

#### Key Terminology Differences

**Schedule vs Job**

* **dbt Cloud**<sup>TM</sup> uses "Job" for scheduled executions
* **Paradime** uses "Bolt Schedule" or simply "Schedule"

**Bolt**

* **Paradime's Bolt** is the orchestration engine (equivalent to dbt Cloud<sup>TM</sup>'s job scheduling system)

**Turbo CI**

* **Turbo CI** is Paradime’s optimized CI/CD system
* Equivalent to dbt Cloud<sup>TM</sup>'s "Continuous Integration Job" / Slim CI but with enhanced features

**Code IDE vs Studio IDE**

* Both are browser-based development environments

**DinoAI**

* **Paradime's AI coding assistant**
* No direct equivalent in dbt Cloud<sup>TM</sup>

**`paradime_schedules.yml`**

* **Paradime-specific** git-tracked YAML file for defining schedules
* dbt Cloud<sup>TM</sup> jobs are primarily configured through UI


# Transferring dbt™️ Jobs

You can migrate all your dbt Cloud™ jobs in Paradime with one click.

{% hint style="warning" %}
Note that access to the dbt Cloud™ API is included only on the **Team** and **Enterprise** plans.
{% endhint %}

### Create Service Token <a href="#id-1-create-service-token" id="id-1-create-service-token"></a>

In dbt Cloud™, navigate to the account setting screen by clicking on the ⚙️ on the top-right of your screen and in the left panel select the Service Token option from the menu.

Click on the `New Token` option to create a new Service Token. Set `paradime` as token name and make sure to set:

* Permission set as `Account Admin`
* Project: select the dbt Cloud™ project that matches the Paradime workspace you are importing into

You can then click on the Save button on the bottom right of the screen to create your Service Token.

{% hint style="info" %}
**A dbt Cloud™ project maps to a Paradime workspace.** The import brings every job the token can access into your **current** Paradime workspace. Scope the token (and the connection) to the single dbt Cloud™ project that corresponds to that workspace so you don't pull another project's jobs into the wrong place. To import several projects, repeat the connection from each matching Paradime workspace.
{% endhint %}

{% hint style="danger" %}
You will not be able to view this token again after generating it, so make sure to save it securely as we will need this in the next step to setup the integration.
{% endhint %}

This will allow Paradime to read the configurations details of your dbt™️ jobs, so that we can import these as Bolt Schedules. You can find more details about dbt Cloud™ API Service Token permissions [here](https://docs.getdbt.com/docs/dbt-cloud-apis/service-tokens).

<figure><img src="/files/ymenbvrX3445f8Zab4br" alt=""><figcaption></figcaption></figure>

### Connect the dbt Cloud™ integration <a href="#id-2-connect-paradime-to-dbt-cloud" id="id-2-connect-paradime-to-dbt-cloud"></a>

The dbt Cloud™ integration can be connected by a Paradime Admin. To enable the integration, in Paradime navigate to your Account Settings > Integrations, find the dbt Cloud™ integration and click *Connect*.

* Host Name: simply copy your dbt Cloud™ url and paste it in this field, make sure to include the account as in the below example
* Service Account Token: enter the Token you have generated in the previous step in dbt Cloud™

Click on Test Connection to validate that Paradime can connect to dbt Cloud™ API and start the import.

<figure><img src="/files/iE4rhojCDCVJc0Q24dFA" alt=""><figcaption></figcaption></figure>

### View Imported dbt Cloud™ Jobs in Paradime <a href="#id-3-view-imported-dbt-cloud-jobs-and-open-a-pull-request" id="id-3-view-imported-dbt-cloud-jobs-and-open-a-pull-request"></a>

After the import is completed, Paradime will create new Bolt schedules using the dbt™️ Cloud Jobs' configurations.

You can now click on **view imported jobs** to view all the imported dbt™️ Cloud jobs in the Paradime Bolt UI. When ready, simply turn them on by updating the cron schedule configurations.

{% hint style="info" %}
**Every imported schedule is created&#x20;*****paused*****.** Nothing runs automatically on import — you review each schedule and enable it from the Bolt UI when you're ready.
{% endhint %}

### How dbt Cloud™ Jobs are mapped to Bolt Schedules <a href="#field-mapping" id="field-mapping"></a>

Each dbt Cloud™ job is translated into a Bolt schedule. The table below shows how the job's configuration is carried across.

| dbt Cloud™ job       | Bolt schedule             | Notes                                                                                                                                          |
| -------------------- | ------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- |
| Job **name**         | Schedule **display name** | The job name is shown as the schedule's name in the UI. A unique slug is generated as the schedule's internal identifier.                      |
| **Commands / steps** | **Commands**              | Copied across exactly as defined.                                                                                                              |
| **Schedule (cron)**  | **Cron**                  | Used when the job runs on a schedule. Jobs that run only on triggers (CI, merge, on-completion) are imported with the schedule set to **OFF**. |
| **Custom branch**    | **Git branch**            | Imported only when the job has a custom branch enabled. Otherwise the schedule uses your Paradime workspace's default branch.                  |
| **Description**      | **Description**           | Copied across.                                                                                                                                 |
| —                    | **Status**                | Always imported **Paused**.                                                                                                                    |
| —                    | **Environment**           | Always set to `production`. The warehouse connection is taken from your Paradime workspace, not from dbt Cloud™.                               |

#### Trigger types

dbt Cloud™ jobs can run in several ways. Each trigger type maps to the equivalent Bolt concept:

| dbt Cloud™ trigger                          | Bolt schedule                                  | Result                                                                     |
| ------------------------------------------- | ---------------------------------------------- | -------------------------------------------------------------------------- |
| Runs **on a schedule** (cron)               | Cron schedule                                  | The cron expression is preserved.                                          |
| Runs **on merge**                           | Trigger on merge                               | The schedule runs when changes are merged.                                 |
| **CI job** that defers to another job       | **Turbo CI**                                   | Runs against the deferred job, schedule set to OFF.                        |
| Job that **defers to another job** (non-CI) | **Deferred schedule**                          | Defers to the referenced schedule.                                         |
| Runs **on completion of another job**       | **Schedule trigger** (run on another schedule) | Chains to the parent schedule; runs after it succeeds/fails as configured. |

### What is *not* imported <a href="#not-imported" id="not-imported"></a>

Some dbt Cloud™ settings have no Bolt equivalent, or are managed elsewhere in Paradime. These are intentionally left at their Paradime default:

| dbt Cloud™ setting                                                                              | Behaviour in Paradime                                                                                                                                       |
| ----------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Environment-level deferral** (a job deferring to an *environment* rather than a specific job) | Not imported — Bolt defers to a *schedule*, not an environment. The schedule is imported without a deferral; set it manually on review (see caveats below). |
| **Threads** / **target name**                                                                   | Taken from your Paradime environment, not the job.                                                                                                          |
| **Timezone**                                                                                    | Schedules run in **UTC**.                                                                                                                                   |
| Generate docs, compare changes, run timeout, run-on-draft-PR, resource class, retries           | No Bolt equivalent — dropped.                                                                                                                               |

### Re-importing jobs <a href="#re-importing" id="re-importing"></a>

{% hint style="warning" %}
**Re-importing creates&#x20;*****duplicate*****&#x20;schedules.** The importer does not yet match against previously-imported jobs, so running the import a second time creates a fresh copy of every job rather than updating the existing schedules.
{% endhint %}

What this means in practice:

* **Changes are not synced.** If you change a job's cron or commands in dbt Cloud™ and re-import, you get a new schedule alongside the old one — the original is not updated.
* **Deletions are not detected.** Jobs removed in dbt Cloud™ remain in Paradime.
* **To avoid duplicates**, delete the previously-imported schedules before re-importing.

### Known caveats & limitations <a href="#caveats" id="caveats"></a>

* **Schedules import paused.** Review and enable each schedule from the Bolt UI when ready.
* **`state:modified+` commands need a deferral target.** Because environment-level deferral isn't imported, a command such as `dbt build --select state:modified+` is preserved but has nothing to compare against. Set the deferred schedule manually before enabling, or the run will fail / select nothing.
* **Duplicate job names.** dbt Cloud™ allows multiple jobs to share a name. These import as separate schedules that display the same name (each has its own unique slug).
* **A dbt Cloud™ project maps to a Paradime workspace.** The import brings the jobs the token can access into your current workspace — it does not split projects into separate workspaces automatically. Scope the token to the project that matches the workspace, and repeat the import from each workspace to bring across multiple projects.
* **On-completion triggers across projects/accounts.** A schedule that runs on the completion of another job is only chained if the parent job is part of the same import. If the parent isn't imported (for example, it lives in another project or account that wasn't part of this import), the schedule is imported paused with no trigger.


# Migrating dbt™ jobs from Github Actions to Paradime Bolt

## Overview

This guide walks you through migrating your dbt™ jobs from GitHub Actions to Paradime's Bolt orchestration platform. Paradime offers a purpose-built solution for dbt™ orchestration with features like deferred runs, smart scheduling, and integrated monitoring.

***

## Part 1: Understanding Your Current GitHub Actions Setup

#### Example GitHub Actions Workflow

Here's a typical GitHub Actions workflow for running dbt™ jobs:

```yaml
# .github/workflows/dbt_production.yml
name: dbt Production Run

on:
  schedule:
    # Run daily at 6 AM UTC
    - cron: '0 6 * * *'
  workflow_dispatch:  # Allow manual triggers
  push:
    branches:
      - main

jobs:
  dbt_run:
    runs-on: ubuntu-latest

    steps:
      - name: Checkout code
        uses: actions/checkout@v3

      - name: Set up Python
        uses: actions/setup-python@v4
        with:
          python-version: '3.11'

      - name: Install dbt
        run: |
          pip install dbt-core dbt-snowflake==1.7.0

      - name: Install dependencies
        run: |
          dbt deps
        env:
          DBT_PROFILES_DIR: .

      - name: Run dbt seed
        run: dbt seed --target prod
        env:
          DBT_SNOWFLAKE_ACCOUNT: ${{ secrets.SNOWFLAKE_ACCOUNT }}
          DBT_SNOWFLAKE_USER: ${{ secrets.SNOWFLAKE_USER }}
          DBT_SNOWFLAKE_PASSWORD: ${{ secrets.SNOWFLAKE_PASSWORD }}
          DBT_SNOWFLAKE_ROLE: ${{ secrets.SNOWFLAKE_ROLE }}
          DBT_SNOWFLAKE_WAREHOUSE: ${{ secrets.SNOWFLAKE_WAREHOUSE }}
          DBT_SNOWFLAKE_DATABASE: ${{ secrets.SNOWFLAKE_DATABASE }}
          DBT_PROFILES_DIR: .

      - name: Run dbt models
        run: dbt run --target prod
        env:
          DBT_SNOWFLAKE_ACCOUNT: ${{ secrets.SNOWFLAKE_ACCOUNT }}
          DBT_SNOWFLAKE_USER: ${{ secrets.SNOWFLAKE_USER }}
          DBT_SNOWFLAKE_PASSWORD: ${{ secrets.SNOWFLAKE_PASSWORD }}
          DBT_SNOWFLAKE_ROLE: ${{ secrets.SNOWFLAKE_ROLE }}
          DBT_SNOWFLAKE_WAREHOUSE: ${{ secrets.SNOWFLAKE_WAREHOUSE }}
          DBT_SNOWFLAKE_DATABASE: ${{ secrets.SNOWFLAKE_DATABASE }}
          DBT_PROFILES_DIR: .

      - name: Run dbt tests
        run: dbt test --target prod
        env:
          DBT_SNOWFLAKE_ACCOUNT: ${{ secrets.SNOWFLAKE_ACCOUNT }}
          DBT_SNOWFLAKE_USER: ${{ secrets.SNOWFLAKE_USER }}
          DBT_SNOWFLAKE_PASSWORD: ${{ secrets.SNOWFLAKE_PASSWORD }}
          DBT_SNOWFLAKE_ROLE: ${{ secrets.SNOWFLAKE_ROLE }}
          DBT_SNOWFLAKE_WAREHOUSE: ${{ secrets.SNOWFLAKE_WAREHOUSE }}
          DBT_SNOWFLAKE_DATABASE: ${{ secrets.SNOWFLAKE_DATABASE }}
          DBT_PROFILES_DIR: .

      - name: Notify on failure
        if: failure()
        run: echo "Job failed - send notification"

```

#### What This Workflow Does

1. **Triggers** on a daily schedule, manual dispatch, or push to main
2. **Installs** dbt™ and dependencies
3. **Runs** dbt™ seed, run, and test commands
4. **Uses** GitHub Secrets for warehouse credentials
5. **Notifies** on failure (basic)

***

## Part 2: Prerequisites for Paradime Migration

Before migrating, ensure you have:

#### 1. Paradime Workspace Setup

* Active Paradime account
* Workspace created and configured
* Access to the Bolt application

#### 2. Data Warehouse Connection

* Production connection configured in Paradime
* This connection should have the same permissions as your GitHub Actions credentials
* Navigate to: **Settings → Connections** to set this up

#### 3. Git Repository Connected

* Your dbt™ project repository connected to Paradime
* Git credentials configured
* Navigate to: **Settings → Git Integration**

#### 4. GitHub App Integration (for Native CI/CD)

* **Paradime GitHub App installed** in your GitHub organization
* This enables native Turbo CI and Continuous Deployment without manual GitHub Actions
* **Installation Guide**: [docs.paradime.io/app-help/documentation/integrations/ci-cd/github](https://docs.paradime.io/app-help/documentation/integrations/ci-cd/github)

#### 5. dbt™ Project

* Your dbt™ project accessible in Paradime IDE
* Models materialized and tested in development

***

## Part 3: Creating Your First Paradime Schedule

### Method 1: Using the Bolt UI (Recommended for Beginners)

#### Step 1: Access Bolt

1. Log into your Paradime workspace
2. Navigate to **Bolt** from the left sidebar
3. Click **Create Schedule**

#### Step 2: Configure Schedule Settings

Fill in the basic settings:

* **Type**: Choose **Standard** (equivalent to a basic GitHub Actions job)
* **Name**: `daily_production_run` (descriptive name for your schedule)
* **Git Branch**: `main` (or your production branch)
* **Owner Email**: Your email address

#### Step 3: Add Commands

Add your dbt™ commands in sequence (equivalent to the steps in your GitHub Action):

```
dbt seed --target prod
dbt run --target prod
dbt test --target prod

```

{% hint style="success" %}
**Pro Tip**: Each command runs sequentially. If one fails, subsequent commands won't execute.
{% endhint %}

#### Step 4: Configure Trigger

Select your trigger type:

* **Scheduled Run**: For time-based execution (like cron)
  * Choose **Cron Schedule**: `0 6 * * *` (daily at 6 AM UTC)
  * Or use presets like `@daily`, `@hourly`, `@weekly`
* **On Merge**: Trigger when PR is merged to specified branch
* **On Run Completion**: Chain jobs together

#### Step 5: Set Up Notifications

Configure alerts to know when jobs succeed or fail:

* **Slack Notifications**:
  * Toggle **Slack Notify On**: Select `failed` and/or `passed`
  * Enter channel: `#data-team-alerts`
* **Email Notifications**:
  * Toggle **Email Notify On**: Select `failed` and/or `passed`
  * Enter email addresses

#### Step 6: Deploy

Click **Create Schedule** to deploy your job.

***

### Method 2: Using Schedules as Code (YAML)

For teams that prefer infrastructure-as-code, Paradime supports YAML-based schedules.

#### Step 1: Create `paradime_schedules.yml`

In your dbt™ project root, create or edit `paradime_schedules.yml`:

```yaml
# paradime_schedules.yml
- name: daily_production_run
  schedule: "0 6 * * *"  # Daily at 6 AM UTC
  environment: production
  git_branch: main
  commands:
    - dbt seed --target prod
    - dbt run --target prod
    - dbt test --target prod
  slack_on:
    - failed
    - passed
  slack_notify:
    - "#data-team-alerts"
  email_on:
    - failed
  email_notify:
    - "data-team@company.com"

```

#### Step 2: Commit and Push

```bash
git add paradime_schedules.yml
git commit -m "Add production schedule"
git push

```

#### Step 3: Sync in Paradime

Paradime will automatically detect and sync your schedule configuration.

***

## Part 4: Migration Mapping Guide

Here's how GitHub Actions concepts map to Paradime Bolt:

| GitHub Actions           | Paradime Bolt             | Notes                         |
| ------------------------ | ------------------------- | ----------------------------- |
| `on.schedule.cron`       | Cron Schedule trigger     | Same cron syntax              |
| `on.push.branches`       | On Merge trigger          | Triggers on merge to branch   |
| `on.workflow_dispatch`   | Manual run button         | Available in Bolt UI          |
| `jobs.steps`             | Commands list             | Sequential execution          |
| `env` secrets            | Connection credentials    | Managed in Paradime settings  |
| Notifications            | Slack/Email notifications | Built into Bolt               |
| `runs-on`                | Paradime environment      | Managed infrastructure        |
| `uses: actions/checkout` | Automatic                 | Paradime handles git checkout |

***

## Part 5: Advanced Features in Paradime

#### Native GitHub Integration Benefits

**With the Paradime GitHub App installed, you get:**

* **Automatic PR checks**: Turbo CI runs automatically on every pull request
* **Merge triggers**: Deploy changes instantly when PRs are merged
* **Lineage Diff comments**: Automated comments showing downstream impact on Looker, Tableau, ThoughtSpot, and dbt™ mesh
* **No GitHub Actions needed**: Paradime handles triggering and status reporting
* **Centralized logs**: All CI/CD runs visible in Bolt dashboard

**Installation**: Follow the guide at [docs.paradime.io/app-help/documentation/integrations/ci-cd/github](https://docs.paradime.io/app-help/documentation/integrations/ci-cd/github)

***

### Deferred Runs (Optimization)

Paradime offers **Deferred Schedules** that only run changed models:

```yaml
- name: incremental_run
  schedule: "@hourly"
  environment: production
  git_branch: main
  deferred_schedule:
    enabled: true
    deferred_schedule_name: daily_production_run
    successful_run_only: true
  commands:
    - dbt run -s state:modified+ --target prod
    - dbt test -s state:modified+ --target prod

```

**Benefits**:

* Faster execution (only modified models)
* Lower compute costs
* Smart state comparison

### Turbo CI for Pull Requests

**🎉 Paradime supports native GitHub integration for Turbo CI!** No need to write GitHub Actions workflows manually.

#### Option 1: Native GitHub App Integration (Recommended)

The easiest way to enable Turbo CI is through Paradime's native GitHub app, which automatically triggers CI checks when you open a pull request.

**Setup Steps:**

1. **Install the Paradime GitHub App**:
   * Navigate to **Settings > Integrations** in Paradime
   * Click **Connect** next to GitHub Integration
   * Follow the authentication flow and select your repositories
   * Click **Install and authorize**
   * Complete the user-level OAuth by going to **Profile > Profile Settings**
2. **Create a Turbo CI Schedule in Bolt**:

   ```yaml
   - name: turbo_ci_run
     schedule: "OFF"  # Only runs on PR
     environment: development
     git_branch: main
     deferred_schedule:
       enabled: true
       deferred_schedule_name: daily_production_run
     commands:
       - dbt build -s state:modified+ --target ci
     slack_on:
       - failed
     slack_notify:
       - "#dev-team"

   ```
3. **That's it!** When you open a pull request, Paradime will automatically:
   * Trigger the Turbo CI schedule
   * Build modified models in a temporary schema (`paradime_turbo_ci_pr_<commit_sha>`)
   * Run tests on changed models
   * Post status checks directly to your PR

**📚 Documentation**: [GitHub Integration Setup](https://docs.paradime.io/app-help/documentation/integrations/ci-cd/github) | [Turbo CI Guide](https://docs.paradime.io/app-help/documentation/bolt/ci-cd/turbo-ci)

#### Option 2: Manual GitHub Actions (Alternative)

If you prefer to manage your own GitHub Actions workflow or need custom logic:

```yaml
# .github/workflows/paradime_turbo_ci.yml
name: Paradime Turbo CI

on:
  pull_request:
    branches: [main]

jobs:
  turbo_ci:
    runs-on: ubuntu-latest
    steps:
      - name: Run Paradime Turbo CI
        run: |
          pip install paradime-io
          paradime bolt run "turbo_ci_run" --branch ${{ github.sha }} --wait
        env:
          PARADIME_API_KEY: ${{ secrets.PARADIME_API_KEY }}
          PARADIME_API_SECRET: ${{ secrets.PARADIME_API_SECRET }}
          PARADIME_API_ENDPOINT: ${{ secrets.PARADIME_API_ENDPOINT }}

```

### Continuous Deployment on Merge

**🎉 Paradime supports native GitHub integration for Continuous Deployment!** Automatically deploy when PRs are merged.

#### Option 1: Native GitHub App Integration (Recommended)

With the Paradime GitHub app installed (see Turbo CI setup above), you can enable automatic deployments on merge without any GitHub Actions configuration.

**Setup in Bolt UI:**

1. Create or edit a schedule in Bolt
2. Set **Schedule Type** to **Deferred**
3. Enable **Trigger on Merge**
4. Select your production branch (e.g., `main`)
5. Configure your commands:

**Or via YAML:**

```yaml
- name: continuous_deployment
  schedule: "OFF"
  environment: production
  git_branch: main
  trigger_on_merge: true
  deferred_schedule:
    enabled: true
    deferred_schedule_name: daily_production_run
    successful_run_only: true
  commands:
    - dbt run -s state:modified+ --target prod
    - dbt test -s state:modified+ --target prod
  slack_on:
    - failed
    - passed
  slack_notify:
    - "#deployment-alerts"

```

**How it works:**

* When a PR is merged to your production branch, Paradime automatically triggers the schedule
* Only modified models are deployed (using state comparison)
* Status updates are posted to Slack
* No GitHub Actions workflow needed!

**📚 Documentation**: [GitHub Native CD Guide](https://docs.paradime.io/app-help/documentation/bolt/ci-cd/continuous-deployment-with-bolt/continuous-deployment)

#### Option 2: Manual GitHub Actions (Alternative)

If you need custom deployment logic or prefer managing workflows yourself:

```yaml
# .github/workflows/paradime_cd.yml
name: Paradime Continuous Deployment

on:
  push:
    branches: [main]

jobs:
  deploy:
    runs-on: ubuntu-latest
    steps:
      - name: Deploy to Production
        run: |
          pip install paradime-io
          paradime bolt run "continuous_deployment" --wait
        env:
          PARADIME_API_KEY: ${{ secrets.PARADIME_API_KEY }}
          PARADIME_API_SECRET: ${{ secrets.PARADIME_API_SECRET }}
          PARADIME_API_ENDPOINT: ${{ secrets.PARADIME_API_ENDPOINT }}

```

***

## Part 6: Step-by-Step Migration Process

#### Phase 1: Parallel Run (Week 1)

1. **Keep GitHub Actions running** as-is
2. **Create equivalent schedules** in Paradime Bolt
3. **Monitor both** for consistency
4. **Compare results** and timing

#### Phase 2: Validation (Week 2)

1. **Verify all jobs** run successfully in Paradime
2. **Test notifications** (Slack/email)
3. **Check monitoring** and logs
4. **Validate data quality** remains consistent

#### Phase 3: Cutover (Week 3)

1. **Disable GitHub Actions** schedules (comment out cron)
2. **Keep GitHub Actions files** for emergency fallback
3. **Monitor Paradime** closely for first few days
4. **Update documentation** and runbooks

#### Phase 4: Cleanup (Week 4+)

1. **Archive GitHub Actions** workflows
2. **Update team documentation**
3. **Train team** on Bolt interface
4. **Optimize schedules** using deferred runs

***

## Part 7: Troubleshooting Common Issues

#### Issue 1: Schedule Not Running

**Check**:

* Schedule is not paused (look for pause icon)
* Cron syntax is correct
* Branch exists and is accessible
* Production connection is active

#### Issue 2: Commands Failing

**Check**:

* Connection credentials are valid
* Target profile exists in `profiles.yml`
* Models exist in specified branch
* Check logs in Bolt → Run History

#### Issue 3: Notifications Not Sending

**Check**:

* Slack workspace is connected (Settings → Integrations)
* Channel names are correct (include `#`)
* Email addresses are valid
* Notification toggles are enabled

#### Issue 4: Git Sync Issues

**Check**:

* Git credentials are valid
* Branch exists
* Repository is accessible
* Try manual sync in Settings

***

## Part 8: Best Practices

#### 1. Naming Conventions

Use descriptive, consistent names:

* ✅ `daily_full_refresh`
* ✅ `hourly_incremental_models`
* ❌ `schedule1`
* ❌ `test`

#### 2. Start Simple

Begin with **Standard schedules**, then adopt:

* Deferred runs for optimization
* Turbo CI for PR validation
* On Run Completion for dependencies

#### 3. Monitor Proactively

* Set up Slack notifications for all production jobs
* Review Run History weekly
* Check SLA compliance in analytics

#### 4. Use Version Control

* Keep `paradime_schedules.yml` in git
* Review changes in PRs
* Document major changes

#### 5. Leverage Deferred Runs

After establishing baseline schedules:

* Identify frequently-modified models
* Create deferred schedules for efficiency
* Monitor compute savings

***

## Part 9: Quick Reference

#### Common Cron Schedules

```
@hourly          # Every hour at minute 0
@daily           # Every day at midnight UTC
@weekly          # Every Sunday at midnight UTC
0 6 * * *        # Daily at 6 AM UTC
0 */4 * * *      # Every 4 hours
0 9 * * 1-5      # Weekdays at 9 AM UTC

```

#### Essential Commands

```bash
# Run all models
dbt run

# Run specific model
dbt run -s my_model

# Run modified models only
dbt run -s state:modified+

# Test all models
dbt test

# Full refresh
dbt run --full-refresh
```

#### Useful Links

* [Paradime Documentation](https://docs.paradime.io/)
* [Bolt Schedules Guide](https://docs.paradime.io/app-help/documentation/bolt)
* [Schedule Types Overview](https://docs.paradime.io/app-help/guides/paradime-101/running-dbt-in-production-with-bolt/understanding-schedule-types-and-triggers)

***

## Conclusion

Migrating from GitHub Actions to Paradime Bolt offers:

* **Simplified management**: No infrastructure to maintain
* **Better observability**: Built-in monitoring and analytics
* **Cost optimization**: Deferred runs and smart scheduling
* **Native dbt™ integration**: Purpose-built for dbt™ workflows
* **Team collaboration**: Centralized scheduling and monitoring

Start with a simple schedule migration, validate in parallel, then gradually adopt advanced features like deferred runs and Turbo CI for maximum efficiency.

***

**Questions or Issues?**

* Check the [Paradime Help Center](https://support.claude.com/)
* Review the [Bolt Documentation](https://docs.paradime.io/app-help/documentation/bolt)
* Contact Paradime support for assistance


# Using dlt (data load tool) Pipelines in Paradime


# dlt Source and Destination Credentials via Environment Variables

Learn how dlt (data load tool) to maps environment variables to source and destination credentials.

## Setting Environment Variables in dlt

In this guide, you'll learn how to configure credentials and secrets for your dlt pipelines using environment variables. By the end of this 5-minute tutorial, you'll understand the naming convention, how to translate your `secrets.toml` keys into environment variable names, and how to apply this across common use cases.

***

### How dlt Reads Configuration

dlt retrieves configuration and secrets from multiple locations in the following order of priority:

1. **Environment Variables** — highest priority; if a value is found here, dlt stops searching
2. **`secrets.toml`** — for secrets like API keys and passwords
3. **`config.toml`** — for non-sensitive configuration values
4. **Default values** — defined in your pipeline code

***

### Naming Convention

Environment variables follow a specific naming convention that maps directly to the structure of your `secrets.toml` or `config.toml` files.

The rules are simple:

* All letters are **capitalized**
* Nested sections (dots `.` in TOML) are replaced with **double underscores `__`**

For example, the following `secrets.toml` entry:

```toml
[destination.bigquery.credentials]
project_id = "my_project"
private_key = "my_key"
client_email = "my_email@project.iam.gserviceaccount.com"
```

Translates to these environment variables:

```bash
DESTINATION__BIGQUERY__CREDENTIALS__PROJECT_ID="my_project"
DESTINATION__BIGQUERY__CREDENTIALS__PRIVATE_KEY="my_key"
DESTINATION__BIGQUERY__CREDENTIALS__CLIENT_EMAIL="my_email@project.iam.gserviceaccount.com"
```

{% hint style="info" %}
**Quick rule:** Replace every `.` with `__` and capitalize everything. `sources.pipedrive.pipedrive_api_key` becomes `SOURCES__PIPEDRIVE__PIPEDRIVE_API_KEY`.
{% endhint %}

***

## Setting Environment Variables

#### On Paradime

In Paradime, navigate to your [environment variable settings](/app-help/documentation/settings/environment-variables) in your workspace and set the key-value pair for your **Code IDE** and **Bolt** sections.

#### On Linux / macOS

Export variables directly in your terminal session:

```bash
export DESTINATION__BIGQUERY__CREDENTIALS__PROJECT_ID="my_project"
export DESTINATION__BIGQUERY__CREDENTIALS__PRIVATE_KEY="my_key"
export DESTINATION__BIGQUERY__CREDENTIALS__CLIENT_EMAIL="my_email@project.iam.gserviceaccount.com"
```

#### On Windows (Command Prompt)

```cmd
set DESTINATION__BIGQUERY__CREDENTIALS__PROJECT_ID=my_project
```

#### In a `.env` File (Local Development)

For local development, create a `.env` file in your project root and use `python-dotenv` to load it automatically:

```bash
# .env
DESTINATION__BIGQUERY__CREDENTIALS__PROJECT_ID="my_project"
DESTINATION__BIGQUERY__CREDENTIALS__PRIVATE_KEY="my_key"
DESTINATION__BIGQUERY__CREDENTIALS__CLIENT_EMAIL="my_email@project.iam.gserviceaccount.com"
```

Then load it in your pipeline script:

```python
from dotenv import load_dotenv
load_dotenv()

import dlt
# Your pipeline code here
```

{% hint style="warning" %} **Never commit your `.env` file to version control.** Add it to your `.gitignore` to keep secrets out of your repository. {% endhint %}

***

### Examples

#### Example 1: Source Credentials (Pipedrive)

The equivalent `secrets.toml` entry:

```toml
[sources.pipedrive]
pipedrive_api_key = "abc123"
```

As an environment variable:

```bash
export SOURCES__PIPEDRIVE__PIPEDRIVE_API_KEY="abc123"
```

***

#### Example 2: Destination Credentials (BigQuery)

The equivalent `secrets.toml` entry:

```toml
[destination.bigquery.credentials]
project_id = "my_project"
private_key = "-----BEGIN RSA PRIVATE KEY-----\n..."
client_email = "service@project.iam.gserviceaccount.com"
```

As environment variables:

```bash
export DESTINATION__BIGQUERY__CREDENTIALS__PROJECT_ID="my_project"
export DESTINATION__BIGQUERY__CREDENTIALS__PRIVATE_KEY="-----BEGIN RSA PRIVATE KEY-----\n..."
export DESTINATION__BIGQUERY__CREDENTIALS__CLIENT_EMAIL="service@project.iam.gserviceaccount.com"
```

***

#### Example 3: Mixing Environment Variables with Existing Variables

If you already have credentials stored under different environment variable names (for example, from another tool), you can map them to the dlt naming convention in your pipeline script:

```python
import dlt
import os

# Map existing environment variables to dlt's expected naming convention
os.environ["DESTINATION__CREDENTIALS__CLIENT_EMAIL"] = os.environ.get("BIGQUERY_CLIENT_EMAIL")
os.environ["DESTINATION__CREDENTIALS__PRIVATE_KEY"] = os.environ.get("BIGQUERY_PRIVATE_KEY")
os.environ["DESTINATION__CREDENTIALS__PROJECT_ID"] = os.environ.get("BIGQUERY_PROJECT_ID")

# For source credentials
dlt.secrets["sources.credentials.client_email"] = os.environ.get("SHEETS_CLIENT_EMAIL")
dlt.secrets["sources.credentials.private_key"] = os.environ.get("SHEETS_PRIVATE_KEY")
dlt.secrets["sources.credentials.project_id"] = os.environ.get("SHEETS_PROJECT_ID")
```

***

#### Example 4: Filesystem Source (AWS S3)

The equivalent `secrets.toml` entry:

```toml
[sources.filesystem.credentials]
aws_access_key_id = "AKIAIOSFODNN7EXAMPLE"
aws_secret_access_key = "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY"
```

As environment variables:

```bash
export SOURCES__FILESYSTEM__CREDENTIALS__AWS_ACCESS_KEY_ID="AKIAIOSFODNN7EXAMPLE"
export SOURCES__FILESYSTEM__CREDENTIALS__AWS_SECRET_ACCESS_KEY="wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY"
```

***

### How dlt Searches for a Value

When a configuration value is missing, dlt logs the exact lookup path it tried. This makes it easy to diagnose issues. For example, if a `password` field is missing for a Postgres destination, dlt will report something like:

```
In Environment Variables key DESTINATION__POSTGRES__CREDENTIALS__PASSWORD was not found.
In Environment Variables key DESTINATION__CREDENTIALS__PASSWORD was not found.
In Environment Variables key CREDENTIALS__PASSWORD was not found.
In secrets.toml key destination.postgres.credentials.password was not found.
```

Use these messages to confirm the exact environment variable name your pipeline expects.

***

### Best Practices

* **Use environment variables in CI/CD and production** to avoid storing secrets in files that could be accidentally committed to version control.
* **Rely on dlt's error messages** when debugging missing credentials — they show the exact keys and providers that were checked.
* **Don't mix providers unnecessarily** — define each credential in one place to avoid confusion over which value dlt is picking up.


# Reuse dbt™ Connection Credentials in dlt Pipelines - BigQuery

Learn how to reuse your dbt™ profiles.yml connection credentials to authenticate dlt pipelines.

In this guide, you'll learn how to reuse your existing BigQuery connection credentials from `profiles.yml` to authenticate dlt pipelines — without duplicating credentials or managing separate config files.

***

### What the Script Expects from profiles.yml

The loader reads your active target directly from `dbt_project.yml` and looks it up in `profiles.yml`. A typical BigQuery profile looks like this:

```yaml
my_project:
  target: prod
  outputs:
    dev:
      type: bigquery
      method: service-account
      project: my-gcp-project
      dataset: my_dataset
      location: US
      keyfile_json:
        type: service_account
        project_id: my-gcp-project
        private_key_id: abc123
        private_key: "-----BEGIN RSA PRIVATE KEY-----\n..."
        client_email: my-sa@my-gcp-project.iam.gserviceaccount.com
```

The script will load whichever target is set as `target:` in your profile — in this case `prod`.

{% hint style="info" %}
**`ingestion` target is optional.**

The `get_credentials_by_environment()` method uses an `ingestion` target as a fallback when your default target is `prod`, to avoid accidentally running pipelines against production. This is only relevant if you call that specific method — the rest of this guide uses `get_active_credentials()`, which works with any target name.
{% endhint %}

***

### 1. Set Up the Credentials Loader

Create a `dlthub` folder at the root of your dbt™ project and add the `profile_connection_credentials.py` script inside it.

```
your-dbt-project/
├── models/
├── dbt_project.yml
├── dlthub/
│   ├── profile_connection_credentials.py   ← add this script here
│   └── my_pipeline.py
```

***

### 2. Use the Credentials in Your dlt Pipeline

In your dlt pipeline file (e.g., `my_pipeline.py`), import the loader and pass the credentials to your BigQuery destination.

```python
import dlt
from profile_connection_credentials import ProfileConnectionCredentialsLoader

# Load credentials from your active dbt™ profile
loader = ProfileConnectionCredentialsLoader()
creds = loader.get_active_credentials()

# Configure your dlt pipeline using the fetched credentials
pipeline = dlt.pipeline(
    pipeline_name="my_pipeline",
    destination=dlt.destinations.bigquery(
        project_id=creds.get("project"),
        credentials=creds.get("keyfile_json"),
        location=creds.get("location", "US"),
    ),
    dataset_name="my_dataset",
)

# Define and run your pipeline as normal
@dlt.resource
def my_data():
    yield [{"id": 1, "value": "hello"}, {"id": 2, "value": "world"}]

load_info = pipeline.run(my_data())
print(load_info)
```

{% hint style="info" %}
**Note**

The loader automatically detects your `dbt_project.yml` and `profiles.yml` paths, as long as `profile_connection_credentials.py` is placed inside a `dlthub` folder at the root of your dbt™ project as shown above. If you place it elsewhere, pass `project_dir` and `profiles_dir` explicitly to `ProfileConnectionCredentialsLoader()`
{% endhint %}

***

### Appendix: Profile Connection Credentials Loader Script

Add the following script as `profile_connection_credentials.py` in your `dlthub` folder.

```python
#!/usr/bin/env python3
"""
Profile Connection Credentials Loader Module

A reusable utility module for fetching connection credentials from dbt's profiles.yml
on demand. Can be imported and used in different Python scripts.

This module directly parses profiles.yml without requiring dbt-core.

Usage:
    from profile_connection_credentials import ProfileConnectionCredentialsLoader
    
    loader = ProfileConnectionCredentialsLoader(
        project_dir="/path/to/dbt/project",
        profiles_dir="/path/to/profiles/dir"
    )
    
    # Get credentials for a specific profile and target
    creds = loader.get_credentials(profile_name="my_profile", target_name="dev")
    
    # Or get the active target credentials (from dbt_project.yml)
    creds = loader.get_active_credentials()
    
    # Access credential attributes
    print(creds['type'])  # e.g., "bigquery"
    print(creds['project_id'])  # database/project specific
"""

import sys
from pathlib import Path
from typing import Optional, Dict, Any

try:
    import yaml
except ImportError:
    print("Error: PyYAML is not installed")
    print("Install it with: pip install pyyaml")
    sys.exit(1)


class ProfileConnectionCredentialsLoader:
    """
    Load and manage connection credentials from profiles.yml by directly parsing YAML.
    
    This class provides a simple interface to fetch credentials on demand
    without requiring dbt-core to be installed.
    """
    
    def __init__(
        self,
        project_dir: Optional[str] = None,
        profiles_dir: Optional[str] = None
    ):
        """
        Initialize the credentials loader.
        
        Args:
            project_dir: Path to the dbt project directory (auto-detected if not provided)
            profiles_dir: Path to the directory containing profiles.yml
                         (auto-detected if not provided)
        
        Raises:
            FileNotFoundError: If project_dir or profiles_dir don't exist
            FileNotFoundError: If dbt_project.yml or profiles.yml not found
        """
        # Auto-detect paths if not provided
        if not project_dir:
            project_dir = self._detect_project_dir()
        if not profiles_dir:
            profiles_dir = self._detect_profiles_dir(project_dir)
        
        self.project_dir = project_dir
        self.profiles_dir = profiles_dir
        self._profiles_data = None
        self._project_data = None
        
        # Validate directories and files
        self._validate_paths()
    
    @staticmethod
    def _detect_project_dir() -> str:
        """
        Auto-detect the dbt project directory.
        
        Looks for dbt_project.yml relative to this script's location.
        Falls back to /workspace/repository/dbt if not found.
        
        Returns:
            Path to the dbt project directory
        """
        # Try to find dbt_project.yml relative to this script
        # This file is at: <repo>/dbt/dlthub/profile_connection_credentials.py
        # So we go up 2 levels to get to <repo>/dbt
        script_file = Path(__file__).resolve()
        script_dir = script_file.parent.parent  # Go up to dbt project root
        
        if (script_dir / "dbt_project.yml").exists():
            return str(script_dir.resolve())
        
        # Fall back to default location
        return "/workspace/repository/dbt"
    
    @staticmethod
    def _detect_profiles_dir(project_dir: str) -> str:
        """
        Auto-detect the profiles directory.
        
        Looks for profiles.yml in the parent directory of the project.
        Falls back to /workspace if not found.
        
        Args:
            project_dir: Path to the dbt project directory
        
        Returns:
            Path to the profiles directory
        """
        # Try to find profiles.yml in the parent directory of project
        # project_dir is typically <repo>/dbt, so parent is <repo>
        project_path = Path(project_dir).resolve()
        potential_profiles_dir = project_path.parent
        if (potential_profiles_dir / "profiles.yml").exists():
            return str(potential_profiles_dir.resolve())
        
        # Fall back to default location
        return "/workspace"
    
    def _validate_paths(self) -> None:
        """Validate that required directories and files exist."""
        project_path = Path(self.project_dir)
        if not project_path.exists():
            raise FileNotFoundError(f"Project directory does not exist: {self.project_dir}")
        
        dbt_project_yml = project_path / 'dbt_project.yml'
        if not dbt_project_yml.exists():
            raise FileNotFoundError(f"dbt_project.yml not found in: {self.project_dir}")
        
        profiles_path = Path(self.profiles_dir)
        if not profiles_path.exists():
            raise FileNotFoundError(f"Profiles directory does not exist: {self.profiles_dir}")
        
        profiles_yml = profiles_path / 'profiles.yml'
        if not profiles_yml.exists():
            raise FileNotFoundError(f"profiles.yml not found in: {self.profiles_dir}")
    
    def _load_profiles_yaml(self) -> Dict[str, Any]:
        """
        Load and parse the profiles.yml file.
        
        Returns:
            Dictionary containing the parsed profiles.yml content
        """
        if self._profiles_data is not None:
            return self._profiles_data
        
        try:
            profiles_path = Path(self.profiles_dir) / 'profiles.yml'
            with open(profiles_path, 'r') as f:
                self._profiles_data = yaml.safe_load(f) or {}
            return self._profiles_data
        except Exception as e:
            raise RuntimeError(f"Failed to load profiles.yml: {e}")
    
    def _load_project_yaml(self) -> Dict[str, Any]:
        """
        Load and parse the dbt_project.yml file.
        
        Returns:
            Dictionary containing the parsed dbt_project.yml content
        """
        if self._project_data is not None:
            return self._project_data
        
        try:
            project_path = Path(self.project_dir) / 'dbt_project.yml'
            with open(project_path, 'r') as f:
                self._project_data = yaml.safe_load(f) or {}
            return self._project_data
        except Exception as e:
            raise RuntimeError(f"Failed to load dbt_project.yml: {e}")
    
    def get_active_credentials(self) -> Dict[str, Any]:
        """
        Get the credentials for the active target (from dbt_project.yml).
        
        Returns:
            Dictionary of credentials for the active target
        
        Raises:
            RuntimeError: If credentials cannot be loaded
        """
        project_data = self._load_project_yaml()
        profiles_data = self._load_profiles_yaml()
        
        # Get profile name from dbt_project.yml
        profile_name = project_data.get('profile')
        if not profile_name:
            raise RuntimeError("No 'profile' specified in dbt_project.yml")
        
        # Get the profile from profiles.yml
        if profile_name not in profiles_data:
            raise RuntimeError(f"Profile '{profile_name}' not found in profiles.yml")
        
        profile = profiles_data[profile_name]
        
        # Get the target name (default or specified)
        target_name = profile.get('target')
        if not target_name:
            raise RuntimeError(f"No 'target' specified for profile '{profile_name}'")
        
        # Get the target credentials
        outputs = profile.get('outputs', {})
        if target_name not in outputs:
            raise RuntimeError(
                f"Target '{target_name}' not found in profile '{profile_name}'"
            )
        
        credentials = outputs[target_name]
        return credentials
    
    def get_credentials(
        self,
        profile_name: Optional[str] = None,
        target_name: Optional[str] = None
    ) -> Dict[str, Any]:
        """
        Get credentials for a specific profile and target.
        
        If profile_name or target_name are not provided, uses the active
        target from dbt_project.yml.
        
        Args:
            profile_name: Name of the profile (optional)
            target_name: Name of the target within the profile (optional)
        
        Returns:
            Dictionary of credentials for the specified target
        
        Raises:
            RuntimeError: If credentials cannot be loaded
        """
        profiles_data = self._load_profiles_yaml()
        
        # If no specific profile/target requested, use active credentials
        if not profile_name and not target_name:
            return self.get_active_credentials()
        
        # Use provided profile_name or get from dbt_project.yml
        if not profile_name:
            project_data = self._load_project_yaml()
            profile_name = project_data.get('profile')
            if not profile_name:
                raise RuntimeError("No 'profile' specified in dbt_project.yml")
        
        # Get the profile from profiles.yml
        if profile_name not in profiles_data:
            raise RuntimeError(f"Profile '{profile_name}' not found in profiles.yml")
        
        profile = profiles_data[profile_name]
        
        # Use provided target_name or get default from profile
        if not target_name:
            target_name = profile.get('target')
            if not target_name:
                raise RuntimeError(f"No 'target' specified for profile '{profile_name}'")
        
        # Get the target credentials
        outputs = profile.get('outputs', {})
        if target_name not in outputs:
            raise RuntimeError(
                f"Target '{target_name}' not found in profile '{profile_name}'"
            )
        
        credentials = outputs[target_name]
        return credentials
    
    def get_credentials_dict(self) -> Dict[str, Any]:
        """
        Get the active credentials as a dictionary.
        
        Returns:
            Dictionary of credential attributes
        
        Raises:
            RuntimeError: If credentials cannot be loaded
        """
        credentials = self.get_active_credentials()
        
        # Filter out private attributes (starting with _)
        creds_dict = {
            key: value for key, value in credentials.items()
            if not key.startswith('_')
        }
        
        return creds_dict
    
    def get_profile_info(self) -> Dict[str, Any]:
        """
        Get information about the loaded profile.
        
        Returns:
            Dictionary with profile metadata
        """
        project_data = self._load_project_yaml()
        profiles_data = self._load_profiles_yaml()
        
        profile_name = project_data.get('profile')
        profile = profiles_data.get(profile_name, {})
        target_name = profile.get('target')
        
        return {
            'project_name': project_data.get('name'),
            'project_root': str(Path(self.project_dir).resolve()),
            'profiles_dir': str(Path(self.profiles_dir).resolve()),
            'profile_name': profile_name,
            'target_name': target_name,
        }
    
    def get_credential_attribute(self, attribute_name: str) -> Any:
        """
        Get a specific attribute from the active credentials.
        
        Args:
            attribute_name: Name of the credential attribute to retrieve
        
        Returns:
            The value of the requested attribute
        
        Raises:
            KeyError: If the attribute doesn't exist
            RuntimeError: If credentials cannot be loaded
        """
        credentials = self.get_active_credentials()
        
        if attribute_name not in credentials:
            raise KeyError(
                f"Credentials dictionary has no key '{attribute_name}'"
            )
        
        return credentials[attribute_name]
    
    def get_credentials_by_environment(self, non_prod_target: str = 'ingestion') -> Dict[str, Any]:
        """
        Get credentials based on the environment (target).
        
        If the default target is 'prod', uses the non-prod target instead.
        Otherwise, uses the default target.
        
        Args:
            non_prod_target: Target name to use if default is 'prod' (default: 'ingestion')
        
        Returns:
            Dictionary of credentials for the selected target
        
        Raises:
            RuntimeError: If credentials cannot be loaded
        """
        profile_info = self.get_profile_info()
        default_target = profile_info.get('target_name')
        
        # If default target is 'prod', use non-prod target instead
        if default_target == 'prod':
            return self.get_credentials(target_name=non_prod_target)
        else:
            # Use the default target
            return self.get_active_credentials()


# Example usage and testing
if __name__ == "__main__":
    import json
    
    if len(sys.argv) < 2:
        print("Usage:")
        print("  python profile_connection_credentials.py <project_dir> [profiles_dir]")
        print("\nExample:")
        print("  python profile_connection_credentials.py /workspace/repository/dbt /workspace")
        sys.exit(1)
    
    project_dir = sys.argv[1]
    profiles_dir = sys.argv[2] if len(sys.argv) > 2 else "/workspace"
    
    try:
        # Initialize the loader
        loader = ProfileConnectionCredentialsLoader(project_dir, profiles_dir)
        
        # Get profile info
        print("\n" + "=" * 70)
        print("PROFILE INFORMATION")
        print("=" * 70)
        profile_info = loader.get_profile_info()
        for key, value in profile_info.items():
            print(f"{key}: {value}")
        
        # Get active credentials
        print("\n" + "=" * 70)
        print("ACTIVE TARGET CREDENTIALS")
        print("=" * 70)
        creds = loader.get_active_credentials()
        print(f"Type: {creds.get('type', 'N/A')}")
        
        # Print all credential attributes
        creds_dict = loader.get_credentials_dict()
        for key, value in creds_dict.items():
            # Redact sensitive information
            if key in ['password', 'token', 'api_key', 'secret', 'private_key']:
                value = "[REDACTED]"
            print(f"{key}: {value}")
        
        print("\n" + "=" * 70 + "\n")
        
    except Exception as e:
        print(f"❌ Error: {e}")
        import traceback
        traceback.print_exc()
        sys.exit(1)
```


# dlt Data Pipeline - Google Sheets to Snowflake

How to ingest data from Google Sheets into Snowflake using dltHub

This guide will walk you through setting up automated data pipelines that extract data from Google Sheets and load it into Snowflake using Paradime and dltHub (a Python data loading tool).

**By the end, you'll be able to:**

* Set up a Python pipeline project alongside your dbt project
* Extract data from Google Sheets automatically
* Load that data into Snowflake
* Run pipelines both in development and production

### Prerequisites

Before starting, make sure you have:

* **Paradime account** with access to a workspace
* **Google Cloud Platform (GCP) account** for Google Sheets API access
* **Snowflake credentials** with permissions to create schemas and tables
* **Google spreadsheet** with some data

### Part 1: Understanding the Project Structure

Your project will look like this:

```
your-project/
├── dbt_project.yml          # Your existing dbt project configuration
├── pyproject.toml            # Python dependency management (we'll create this)
├── python_dlt/               # Folder for your Python pipelines
│   ├── gsheet_pipeline.py    # Your pipeline script (we'll create this)
│   └── ...                   # dltHub scaffoldings (auto-generated)
└── models/                   # Your existing dbt models
```

### Part 2: Set Up Google Sheets API Access

#### Why do we need this?

To read data from Google Sheets programmatically, you need API credentials from Google Cloud Platform.

#### Steps:

1. **Go to Google Cloud Console**
   * Visit [console.cloud.google.com](https://console.cloud.google.com/)
   * Sign in with your Google account
2. **Create a Service Account** (if you don't have one)
   * In the left menu, go to **IAM & Admin** → **Service Accounts**
   * Click **+ CREATE SERVICE ACCOUNT**
   * Give it a name like "paradime-sheets-reader"
   * Click **CREATE AND CONTINUE**
   * Skip the optional steps and click **DONE**
3. **Enable Google Sheets API**
   * In the search bar at the top, type "Google Sheets API"
   * Click on it and press **ENABLE**
4. **Create API Credentials**
   * Go back to **IAM & Admin** → **Service Accounts**
   * Find your service account and click the three dots (⋮) under **Actions**
   * Select **Manage Keys**
   * Click **ADD KEY** → **Create new key**
   * Choose **JSON** format
   * Click **CREATE** - a JSON file will download automatically
5. **Share Your Google Sheet**
   * Open the JSON file you just downloaded
   * Find the `client_email` field (looks like: `your-service@project.iam.gserviceaccount.com`)
   * Copy this email address
   * Go to your Google Sheet and click **Share**
   * Paste the service account email and give it **Viewer** access
6. **Extract Credentials from JSON**

   Open the downloaded JSON file. You'll need these four values:

   ```json
   {
     "project_id": "your-project-12345",
     "private_key": "-----BEGIN PRIVATE KEY-----\nMIIEvQIBA...",
     "client_email": "your-service@project.iam.gserviceaccount.com",
     "token_uri": "https://oauth2.googleapis.com/token"
   }
   ```
7. **Encode the Private Key**

   The private key needs to be encoded in Base64 format.

   **On Mac/Linux:**

   ```bash
   echo -n 'YOUR_PRIVATE_KEY_HERE' | base64
   ```
8. **On Windows (PowerShell):**

   ```powershell
   [Convert]::ToBase64String([Text.Encoding]::UTF8.GetBytes('YOUR_PRIVATE_KEY_HERE'))
   ```
9. Save this encoded value - you'll use it as `B64_BIGQUERY_PRIVATE_KEY`.

***

### Part 3: Prepare Snowflake Credentials

#### Why RSA Key Authentication?

For security, we use key-pair authentication instead of passwords for automated pipelines.

#### Steps:

**Generate RSA Key Pair** (if you don't have one)

```bash
# Generate encrypted private key
openssl genrsa 2048 | openssl pkcs8 -topk8 -v2 des3 -inform PEM -out rsa_key.p8

# Generate public key
openssl rsa -in rsa_key.p8 -pubout -out rsa_key.pub
```

**Register Public Key in Snowflake**

```sql
-- Connect to Snowflake and run:
ALTER USER your_username SET RSA_PUBLIC_KEY='MIIBIjANBgkqhki...';
```

To get the key value, open `rsa_key.pub` and copy everything between the header/footer lines.

**Format Private Key for Paradime**

Open `rsa_key.p8` in a text editor:

**Original:**

```
-----BEGIN ENCRYPTED PRIVATE KEY-----
MIIFHDBOBgkqhkiG9w0BAQEFAAOCAg8AMIICCgKCAgEA
zdLQw8fH9QxPxHFvqKVhH3zqxKgHHGxXrKJH8fH9QxPx
...more lines...
qwertyuiop
-----END ENCRYPTED PRIVATE KEY-----
```

**Formatted (remove headers and join into one line):**

```
MIIFHDBOBgkqhkiG9w0BAQEFAAOCAg8AMIICCgKCAgEAzdLQw8fH9QxPxHFvqKVhH3zqxKgHHGxXrKJH8fH9QxPx...qwertyuiop
```

**Gather Snowflake Connection Details**

You'll need:

* **Account identifier**: Found in your Snowflake URL (e.g., `xy12345.us-east-1`)
* **Database name**: The database where data will be loaded
* **Warehouse name**: The compute warehouse to use
* **Role**: Your Snowflake role (e.g., `ACCOUNTADMIN`, `SYSADMIN`)
* **Username**: Your Snowflake username
* **Passphrase**: \[optional] if set when generating the private key

***

### Part 4: Configure Environment Variables in Paradime

Environment variables are secure ways to store credentials without hardcoding them in your scripts.

#### Where to Set Them:

You need to set these in **TWO places** in Paradime:

1. [**Code IDE**](/app-help/documentation/settings/environment-variables/code-ide-env-variables) (for development)
   * Click Settings → Environment Variables
   * Scroll to " Code IDE"
2. [**Bolt Scheduler**](/app-help/documentation/settings/environment-variables/bolt-schedule-env-variables) (for production runs)
   * Click Settings → Environment Variables
   * Scroll to " Bolt"

#### Variables to Set:

**Google Sheets Credentials**

| Variable Name                                       | Value                                    | Example                                   |
| --------------------------------------------------- | ---------------------------------------- | ----------------------------------------- |
| `SOURCES__GOOGLE_SHEETS__CREDENTIALS__PROJECT_ID`   | From JSON: `project_id`                  | `my-project-12345`                        |
| `SOURCES__GOOGLE_SHEETS__CREDENTIALS__CLIENT_EMAIL` | From JSON: `client_email`                | `service@project.iam.gserviceaccount.com` |
| `SOURCES__GOOGLE_SHEETS__CREDENTIALS__TOKEN_URI`    | From JSON: `token_uri`                   | `https://oauth2.googleapis.com/token`     |
| `B64_BIGQUERY_PRIVATE_KEY`                          | Base64-encoded private key from Step 2.7 | `LS0tLS1CRUdJTi...`                       |

**Snowflake Credentials**

| Variable Name                                                 | Value                                 | Example                |
| ------------------------------------------------------------- | ------------------------------------- | ---------------------- |
| `DESTINATION__SNOWFLAKE__CREDENTIALS__HOST`                   | Snowflake account identifier          | `xy12345.us-east-1`    |
| `DESTINATION__SNOWFLAKE__CREDENTIALS__DATABASE`               | Target database                       | `ANALYTICS`            |
| `DESTINATION__SNOWFLAKE__CREDENTIALS__WAREHOUSE`              | Warehouse name                        | `COMPUTE_WH`           |
| `DESTINATION__SNOWFLAKE__CREDENTIALS__ROLE`                   | Your role                             | `TRANSFORMER`          |
| `DESTINATION__SNOWFLAKE__CREDENTIALS__USERNAME`               | Your username                         | `john.doe@company.com` |
| `DESTINATION__SNOWFLAKE__CREDENTIALS__PRIVATE_KEY`            | Formatted private key (from Step 3.3) | `MIIFHDBOBgkq...`      |
| `DESTINATION__SNOWFLAKE__CREDENTIALS__PRIVATE_KEY_PASSPHRASE` | Passphrase you created                | `MySecurePass123!`     |

**Development Variable**

| Variable Name       | Value                       | Example |
| ------------------- | --------------------------- | ------- |
| `DEV_SCHEMA_PREFIX` | Your initials or identifier | `JD`    |

{% hint style="info" %}
**Note:** Only set `DEV_SCHEMA_PREFIX` in Code IDE settings, NOT in Bolt. This ensures development data is isolated from production.
{% endhint %}

***

### Part 5: Initialize Your Python Project

Now we'll set up the project structure and dependencies.

#### Step 5.1: Create `pyproject.toml`

In Paradime's Code IDE, at the root level (same folder as `dbt_project.yml`), create a new file called `pyproject.toml`:

```toml
[tool.poetry]
name = "python-pipelines"
version = "0.1.0"
description = "Data pipelines using dltHub"
authors = ["Your Name <your.email@company.com>"]
readme = "README.md"
package-mode = false

[tool.poetry.dependencies]
python = ">=3.11,<3.13"
google-api-python-client = "^2.118.0"
dlt = {extras = ["snowflake"], version = "1.5.0"}

[build-system]
requires = ["poetry-core"]
build-backend = "poetry.core.masonry.api"
```

**What this does:**

* **Poetry** is a Python dependency manager (like npm for JavaScript)
* We're specifying we need Python 3.11 or 3.12
* We're installing `dlt` with Snowflake support and Google API client

#### Step 5.2: Install Dependencies

Open the terminal in Paradime's Code IDE and run:

```bash
poetry install
```

This will:

* Create a virtual environment
* Install dltHub and all required packages
* Generate a `poetry.lock` file (commit this to git!)

#### Step 5.3: Initialize dltHub

Create the `python_dlt` folder and initialize the Google Sheets pipeline:

```bash
mkdir python_dlt
cd python_dlt
dlt init google_sheets snowflake
```

**What this command does:**

* Creates `.dlt/` folder with configuration files
* Downloads the Google Sheets source code
* Creates example pipeline script
* Adds a `.gitignore` file

**Note:** We're using environment variables instead of `.dlt/secrets.toml` for security, so you can ignore that file.

***

### Part 6: Create Your Pipeline Script

Replace the example `google_sheets_pipeline.py` or create a new file called `gsheet_pipeline.py`:

```python
import os
import dlt
import base64
from google_sheets import google_spreadsheet
from typing import List, Sequence

def setup_credentials():
    """
    Decode and set up Google Sheets API credentials.
    
    This function retrieves the base64-encoded private key from environment
    variables, decodes it, and sets it in the format dltHub expects.
    """
    b64_private_key = os.getenv("B64_BIGQUERY_PRIVATE_KEY")
    if not b64_private_key:
        raise ValueError("Missing B64_BIGQUERY_PRIVATE_KEY environment variable")

    try:
        # Decode from base64
        private_key = base64.b64decode(b64_private_key).decode('utf-8')
        # Replace escaped newlines with actual newlines
        private_key = private_key.replace('\\n', '\n').strip()
        # Set for dltHub to use
        os.environ["SOURCES__GOOGLE_SHEETS__CREDENTIALS__PRIVATE_KEY"] = private_key
    except Exception as e:
        raise ValueError(f"Failed to decode private key: {str(e)}")

def get_dataset_name() -> str:
    """
    Determine the Snowflake schema name based on the environment.
    
    Production (scheduled runs): Uses GOOGLE_SHEETS_LOAD
    Development (manual runs): Uses {DEV_SCHEMA_PREFIX}_GOOGLE_SHEETS_LOAD
    
    Returns:
        str: The schema/dataset name to use
    
    Raises:
        ValueError: If neither environment variable is set
    """
    # Paradime automatically sets this when running scheduled jobs
    schedule_run_id = os.getenv("PARADIME_SCHEDULE_RUN_ID")
    # You manually set this in Code IDE settings
    dev_prefix = os.getenv("DEV_SCHEMA_PREFIX")
    
    if schedule_run_id:
        # We're in a scheduled production run
        return "GOOGLE_SHEETS_LOAD"
    elif dev_prefix:
        # We're in development mode
        return f"{dev_prefix}_GOOGLE_SHEETS_LOAD"
    else:
        raise ValueError(
            "Must set either PARADIME_SCHEDULE_RUN_ID (for production) "
            "or DEV_SCHEMA_PREFIX (for development) environment variable"
        )

def load_pipeline(spreadsheet_url_or_id: str, range_names: Sequence[str]) -> None:
    """
    Extract data from Google Sheets and load into Snowflake.
    
    Args:
        spreadsheet_url_or_id: The Google Sheet ID or full URL
                               Example: "1U3NQQrMgodDV8t-OPrHpf9wVgrXU2HKbrJK2effbUjA"
        range_names: List of sheet names or ranges to extract
                    Example: ["Sheet1", "Sheet2!A1:B10"]
    """
    # Set up Google credentials
    setup_credentials()
    
    # Determine which schema to use
    dataset_name = get_dataset_name()
    
    # Create the dltHub pipeline
    pipeline = dlt.pipeline(
        pipeline_name="google_sheets_pipeline",  # Identifier for this pipeline
        destination='snowflake',                 # Where to load data
        dev_mode=False,                          # Disable dev mode (use for testing)
        dataset_name=dataset_name                # Target schema name
    )
    
    # Configure what to extract from Google Sheets
    data = google_spreadsheet(
        spreadsheet_url_or_id=spreadsheet_url_or_id,
        range_names=range_names,           # Only load these specific sheets/ranges
        get_sheets=False,                  # Don't automatically load all sheets
        get_named_ranges=False             # Don't load named ranges
    )
    
    # Run the pipeline
    print(f"Loading data into dataset: {dataset_name}")
    info = pipeline.run(data)
    print(info)

if __name__ == "__main__":
    # CONFIGURE YOUR GOOGLE SHEET HERE
    
    # Option 1: Use the full URL
    spreadsheet_url_or_id = "https://docs.google.com/spreadsheets/d/1U3NQQrMgodDV8t-OPrHpf9wVgrXU2HKbrJK2effbUjA/edit"
    
    # Option 2: Or just the sheet ID (the long string in the URL)
    # spreadsheet_url_or_id = "1U3NQQrMgodDV8t-OPrHpf9wVgrXU2HKbrJK2effbUjA"
    
    # Specify which sheets or ranges to load
    range_names = ["WorldCupMatches"]  # Sheet name(s)
    # range_names = ["Sheet1!A1:D100"]  # Or specific range(s)
    # range_names = ["Sheet1", "Sheet2", "DataSheet!A:Z"]  # Multiple sheets
    
    # Run the pipeline
    load_pipeline(spreadsheet_url_or_id=spreadsheet_url_or_id, range_names=range_names)
```

#### Understanding the Code

**Function Breakdown:**

1. **`setup_credentials()`**: Handles the Google API authentication
   * Retrieves the encoded private key
   * Decodes it from Base64
   * Sets it in the format dltHub expects
2. **`get_dataset_name()`**: Smart schema naming
   * In development: Creates `{YOUR_INITIALS}_GOOGLE_SHEETS_LOAD`
   * In production: Creates `GOOGLE_SHEETS_LOAD`
   * Prevents dev data from mixing with prod data
3. **`load_pipeline()`**: The main extraction and loading logic
   * Sets up credentials
   * Creates a dltHub pipeline object
   * Configures the Google Sheets source
   * Runs the extraction and loading
4. **`if __name__ == "__main__"`**: Configuration section
   * This is where YOU specify which Google Sheet to load
   * And which specific sheets/ranges to extract

***

### Part 7: Run Your First Pipeline

#### Development Run (Testing)

1. **Open Paradime's Code IDE terminal**
2. **Navigate to the pipeline folder:**

   ```bash
   cd python_dlt
   ```
3. **Run the pipeline using Poetry:**

   ```bash
   poetry run python gsheet_pipeline.py
   ```

**What happens:**

* The script executes in the Poetry virtual environment
* Data is extracted from your Google Sheet
* Tables are created in Snowflake under `{YOUR_PREFIX}_GOOGLE_SHEETS_LOAD` schema
* You'll see progress logs in the terminal

**Expected Output:**

```
Loading data into dataset: JD_GOOGLE_SHEETS_LOAD
Pipeline google_sheets_pipeline load step completed in X.XX seconds
1 load package(s) were loaded
```

#### Production Run (Scheduled)

1. **In Paradime, go to Bolt → Schedules**
2. [**Create a new schedule:**](/app-help/documentation/bolt/creating-schedules) name: "Load Google Sheets Data"
3. **Add Commands:**

```
# Install dependencies with Poetry
poetry install

# Run your Python script
poetry run python python_dlt/gsheet_pipeline.py
```

4. **Schedule:** Choose frequency (e.g., "Daily at 6 AM")
5. **Save and enable the schedule**

<figure><img src="/files/2md3d0NwdwWLpwnY3nPM" alt=""><figcaption></figcaption></figure>

**What happens:**

* Paradime automatically sets `PARADIME_SCHEDULE_RUN_ID`
* Data loads into `GOOGLE_SHEETS_LOAD` schema (without your prefix)
* Your dbt models can now reference this data
* Runs automatically on your chosen schedule

***

### Part 8: Using the Loaded Data in dbt

Now that data is in Snowflake, you can create dbt models to transform it:

```sql
-- models/staging/stg_world_cup_matches.sql

with source as (
    select * from {{ source('google_sheets', 'worldcupmatches') }}
),

renamed as (
    select
        year,
        datetime,
        stage,
        stadium,
        city,
        home_team_name,
        home_team_goals,
        away_team_goals,
        away_team_name,
        win_conditions,
        attendance,
        half_time_home_goals,
        half_time_away_goals,
        referee,
        assistant_1,
        assistant_2,
        _dlt_load_id,
        _dlt_id
    from source
)

select * from renamed
```

**Define the source in `sources.yml`:**

```yaml
version: 2

sources:
  - name: google_sheets
    schema: google_sheets_load  # Production schema
    tables:
      - name: worldcupmatches
        description: "Raw World Cup match data from Google Sheets"
```

***

### Troubleshooting Common Issues

<details>

<summary>"Missing B64_BIGQUERY_PRIVATE_KEY environment variable"</summary>

**Cause:** Environment variable not set or not accessible

**Fix:**

* Double-check spelling in Paradime settings
* Ensure you saved the environment variable
* Restart the Code IDE to pick up new variables

</details>

<details>

<summary>"Authentication failed" for Google Sheets</summary>

**Cause:** Service account doesn't have access to the sheet

**Fix:**

1. Open your Google Sheet
2. Click "Share"
3. Add the service account email (from `client_email` in JSON)
4. Give at least "Viewer" permission

</details>

<details>

<summary>"Unable to connect to Snowflake"</summary>

**Cause:** Incorrect Snowflake credentials or network issues

**Fix:**

* Verify all Snowflake environment variables are set correctly
* Test the private key format (no headers, single line)
* Check your Snowflake role has CREATE SCHEMA permissions
* Verify the warehouse is running

</details>

<details>

<summary>Tables created but empty</summary>

**Cause:** Wrong sheet name or range specified

**Fix:**

* Double-check the sheet name spelling (case-sensitive)
* Verify the sheet exists in the Google Sheet
* Try using just the sheet name without range specifiers

</details>

<details>

<summary>"Command not found: poetry"</summary>

**Cause:** Poetry not installed in the environment

**Fix:**

```bash
pip install poetry
```

</details>

### Best Practices

#### 1. Schema Organization

**Development:**

* Use personal prefixes (`JD_`, `SARAH_`, etc.)
* Prevents conflicts when multiple developers test
* Easy to identify and clean up dev data

**Production:**

* Use clean, standard names (`GOOGLE_SHEETS_LOAD`)
* Apply proper access controls
* Document schema purposes

#### 2. Pipeline Configuration

**Keep it flexible:**

```python
# Good - easy to modify
SPREADSHEET_ID = "1U3NQQ..."
SHEET_NAMES = ["Sales2024", "Sales2023"]

# Then use in function call
load_pipeline(SPREADSHEET_ID, SHEET_NAMES)
```

**Version control your changes:**

* Commit `pyproject.toml` and `poetry.lock`
* Track pipeline scripts in git
* Document any configuration changes

#### 3. Error Handling

**Add try-except blocks for robustness:**

```python
try:
    load_pipeline(spreadsheet_url_or_id, range_names)
except Exception as e:
    print(f"Pipeline failed: {str(e)}")
    # Send alert, log to monitoring system, etc.
    raise
```

#### 4. Monitoring

**Things to track:**

* Pipeline run duration
* Number of rows loaded
* Data freshness in Snowflake
* Failed run alerts

**Example logging:**

```python
import logging

logging.basicConfig(level=logging.INFO)
logger = logging.getLogger(__name__)

logger.info(f"Starting pipeline for {spreadsheet_url_or_id}")
info = pipeline.run(data)
logger.info(f"Loaded {info.dataset_name} successfully")
```

#### 5. Incremental Loading

For large sheets, configure incremental loading to only extract new data:

```python
# In your pipeline configuration
data = google_spreadsheet(
    spreadsheet_url_or_id=spreadsheet_url_or_id,
    range_names=range_names,
    get_sheets=False
)

# Add incremental configuration
data = data.add_incremental(
    "timestamp_column",  # Column to use for incrementality
    primary_key="id"     # Unique identifier column
)
```

***

### Next Steps

Now that you have a working pipeline:

1. **Add More Data Sources**
   * Check [dltHub documentation](https://dlthub.com/docs/) for other verified sources
   * Initialize additional pipelines: `dlt init [source] snowflake`
2. **Build Transformations**
   * Create dbt™. models that use your loaded data
   * Add tests and documentation
   * Build downstream analytics models
3. **Automate Everything**
   * Schedule your pipeline in Bolt
   * Set up dbt™ to run after data loads
   * Create alerts for pipeline failures
4. **Explore Advanced Features**
   * Incremental loading strategies
   * Schema evolution handling
   * Custom data transformations in Python
   * Multiple destination support

***

### Additional Resources

* **dltHub Documentation**: [dlthub.com/docs](https://dlthub.com/docs/)
* **dltHub Verified Sources**: [dlthub.com/docs/verified-sources](https://dlthub.com/docs/dlt-ecosystem/verified-sources/)
* **Google Sheets API**: [developers.google.com/sheets](https://developers.google.com/sheets/api)
* **Poetry Documentation**: [python-poetry.org/docs](https://python-poetry.org/docs/)


# PII Anonymization with dbt™ Mesh Setup

## Overview

This document demonstrates how to set up a dbt mesh architecture using Paradime where a parent repository contains PII (Personally Identifiable Information) models, and a child dbt project consumes anonymized subsets of these models.

## Architecture

```
Parent Repo (customer-data-platform)
├── PII Models (private)
├── Anonymized Models (public via mesh)
└── Data transformations

Child Repo (analytics-workspace)
├── Consumes anonymized models from parent
├── Creates analytics models
└── Business intelligence layer

```

## Parent Repository Setup

#### 1. Project Structure

```yaml
# dbt_project.yml (Parent)
name: 'customer_data_platform'
version: '1.0.0'
config-version: 2

model-paths: ["models"]
analysis-paths: ["analyses"]
test-paths: ["tests"]
seed-paths: ["seeds"]
macro-paths: ["macros"]
snapshot-paths: ["snapshots"]

models:
  customer_data_platform:
    # Private PII models - not exposed
    staging:
      +materialized: table
      +group: private_data

    # Public anonymized models - exposed via mesh
    marts:
      anonymized:
        +materialized: table
        +group: public_analytics
        +access: public

```

#### 2. Model Groups Configuration

```yaml
# models/_groups.yml (Parent)
version: 2

groups:
  - name: private_data
    description: "Internal PII and sensitive customer data - not exposed via mesh"
    owner:
      name: "Data Platform Team"
      email: "data-platform@company.com"

  - name: public_analytics
    description: "Anonymized and aggregated data safe for analytics consumption"
    owner:
      name: "Analytics Team"
      email: "analytics@company.com"

```

#### 3. Private PII Models

```sql
-- models/staging/stg_customers.sql (Private - PII)
{{ config(
    materialized='table',
    group='private_data'
) }}

select
    customer_id,
    first_name,           -- PII
    last_name,            -- PII
    email,                -- PII
    phone,                -- PII
    date_of_birth,        -- PII
    created_at,
    updated_at
from {{ source('raw_data', 'customers') }}

```

```sql
-- models/staging/stg_orders.sql (Private - contains PII references)
{{ config(
    materialized='table',
    group='private_data'
) }}

select
    order_id,
    customer_id,          -- Links to PII
    order_date,
    total_amount,
    shipping_address,     -- PII
    billing_address,      -- PII
    created_at
from {{ source('raw_data', 'orders') }}

```

#### 4. Public Anonymized Models (Exposed via Mesh)

```sql
-- models/marts/anonymized/customers_anonymized.sql
{{ config(
    materialized='table',
    group='public_analytics',
    access='public'
) }}

select
    {{ dbt_utils.generate_surrogate_key(['customer_id', 'created_at']) }} as customer_key,
    -- Anonymize age instead of DOB
    case
        when date_diff('year', date_of_birth, current_date()) < 18 then 'Under 18'
        when date_diff('year', date_of_birth, current_date()) between 18 and 25 then '18-25'
        when date_diff('year', date_of_birth, current_date()) between 26 and 35 then '26-35'
        when date_diff('year', date_of_birth, current_date()) between 36 and 50 then '36-50'
        else 'Over 50'
    end as age_group,

    -- Geographic aggregation
    regexp_extract(shipping_address, r'([A-Z]{2})\\\\s+\\\\d{5}') as state_code,

    -- Temporal data (safe to expose)
    date_trunc('month', created_at) as signup_month,
    date_trunc('quarter', created_at) as signup_quarter,

    -- Customer lifecycle info
    case
        when date_diff('day', created_at, current_date()) <= 30 then 'New'
        when date_diff('day', created_at, current_date()) <= 365 then 'Active'
        else 'Established'
    end as customer_segment

from {{ ref('stg_customers') }}

```

```sql
-- models/marts/anonymized/order_metrics.sql
{{ config(
    materialized='table',
    group='public_analytics',
    access='public'
) }}

select
    {{ dbt_utils.generate_surrogate_key(['customer_id', 'order_date']) }} as customer_day_key,
    date_trunc('day', order_date) as order_date,
    date_trunc('month', order_date) as order_month,
    date_trunc('quarter', order_date) as order_quarter,

    -- Aggregated metrics (no PII)
    count(*) as order_count,
    sum(total_amount) as total_revenue,
    avg(total_amount) as avg_order_value,
    min(total_amount) as min_order_value,
    max(total_amount) as max_order_value,

    -- Geographic aggregation
    regexp_extract(shipping_address, r'([A-Z]{2})\\\\s+\\\\d{5}') as state_code

from {{ ref('stg_orders') }}
group by 1, 2, 3, 4, 8

```

## Child Repository Setup

#### 1. Paradime Mesh Dependencies Configuration

```yaml
# dbt_loom.config.yml (Child)
manifests:
  - name: customer_data_platform  # The name of your "producer" project
    type: paradime
    config:
      schedule_name: daily_production_run  # The name of the Bolt schedule in your "producer" project
      # Environment variables for API credentials from the "producer" project
      api_key: ${CUSTOMER_PLATFORM_API_KEY}
      api_secret: ${CUSTOMER_PLATFORM_API_SECRET}
      api_endpoint: ${CUSTOMER_PLATFORM_API_ENDPOINT}

```

#### 2. Project Configuration

```yaml
# dbt_project.yml (Child)
name: 'analytics_workspace'
version: '1.0.0'
config-version: 2

model-paths: ["models"]
analysis-paths: ["analyses"]
test-paths: ["tests"]
seed-paths: ["seeds"]
macro-paths: ["macros"]

models:
  analytics_workspace:
    marts:
      +materialized: table

    reports:
      +materialized: view

```

#### 3. Consuming Parent Models

```sql
-- models/marts/customer_analytics.sql (Child)
{{ config(materialized='table') }}

with customer_base as (
    -- Reference producer project models using two-argument ref
    select * from {{ ref('customer_data_platform', 'customers_anonymized') }}
),

order_summary as (
    select
        customer_day_key,
        order_month,
        order_quarter,
        state_code,
        sum(total_revenue) as monthly_revenue,
        sum(order_count) as monthly_orders,
        avg(avg_order_value) as avg_monthly_aov
    from {{ ref('customer_data_platform', 'order_metrics') }}
    group by 1, 2, 3, 4
)

select
    c.customer_key,
    c.age_group,
    c.customer_segment,
    c.state_code,
    c.signup_quarter,

    -- Aggregated metrics from parent
    coalesce(sum(o.monthly_revenue), 0) as lifetime_value,
    coalesce(sum(o.monthly_orders), 0) as total_orders,
    coalesce(avg(o.avg_monthly_aov), 0) as average_order_value,

    -- Time-based metrics
    date_diff('month', c.signup_month, current_date()) as customer_tenure_months

from customer_base c
left join order_summary o
    on c.customer_key = o.customer_day_key
group by 1, 2, 3, 4, 5, 9

```

#### 4. Business Intelligence Models

```sql
-- models/reports/customer_segmentation_report.sql (Child)
{{ config(materialized='view') }}

select
    age_group,
    customer_segment,
    state_code,

    count(*) as customer_count,
    sum(lifetime_value) as segment_revenue,
    avg(lifetime_value) as avg_lifetime_value,
    avg(total_orders) as avg_orders_per_customer,
    avg(average_order_value) as avg_order_value,
    avg(customer_tenure_months) as avg_tenure_months,

    -- Performance metrics
    sum(lifetime_value) / count(*) as revenue_per_customer,
    sum(total_orders) / count(*) as orders_per_customer

from {{ ref('customer_analytics') }}
group by 1, 2, 3
order by segment_revenue desc

```

## Paradime Configuration

#### 1. Producer Project Setup

**Prerequisites:**

* dbt version 1.7 or greater in both projects
* At least one successful Bolt schedule run in the producer project
* Models with `access: public` configuration

**Producer project requirements:**

> **Ensure you have a Bolt schedule running (e.g., daily\_production\_run)** **This is required for Paradime to fetch model metadata**

#### 2. Consumer Project API Credentials Setup

**Step 1: Generate API credentials in the producer project**

* Navigate to the producer project (`customer_data_platform`)
* Go to Settings → API Keys
* Generate API credentials with "Bolt schedules metadata viewer" capability
* Note down: API Key, API Secret, and API Endpoint

**Step 2: Set Workspace-level Environment Variables (for Bolt schedules)** In the consumer project workspace settings, add:

```
CUSTOMER_PLATFORM_API_KEY=<your_producer_api_key>
CUSTOMER_PLATFORM_API_SECRET=<your_producer_api_secret>
CUSTOMER_PLATFORM_API_ENDPOINT=<your_producer_api_endpoint>

```

**Step 3: Set User-level Environment Variables (for Code IDE)** Each developer in the consumer project must set the same environment variables in their Code IDE settings:

```
CUSTOMER_PLATFORM_API_KEY=<your_producer_api_key>
CUSTOMER_PLATFORM_API_SECRET=<your_producer_api_secret>
CUSTOMER_PLATFORM_API_ENDPOINT=<your_producer_api_endpoint>

```

#### 3. Model Referencing in Consumer Project

Always use the two-argument `ref` function when referencing models from the producer project:

```sql
-- ✅ Correct way to reference producer models
select * from {{ ref('customer_data_platform', 'customers_anonymized') }}

-- ❌ This won't work for cross-project references
select * from {{ ref('customers_anonymized') }}

```

## Security Considerations

#### Access Control

* PII models are in `private_data` group with no public access
* Only anonymized models in `public_analytics` group are exposed
* Child projects can only access explicitly exposed models

### Testing Strategy

#### 1. Parent Project Tests

```sql
-- tests/assert_no_pii_in_public_models.sql
-- Ensure no PII leaks into public models

select *
from {{ ref('customers_anonymized') }}
where customer_key in (
    select customer_id::string
    from {{ ref('stg_customers') }}
)

```

#### 2. Child Project Tests

```sql
-- tests/validate_mesh_data_quality.sql
-- Ensure mesh data meets quality standards

select *
from {{ ref('customer_analytics') }}
where lifetime_value < 0
   or total_orders < 0
   or customer_tenure_months < 0

```

## Best Practices

1. **Regular Security Audits**: Review anonymized models quarterly
2. **Change Management**: Use PR reviews for any changes to public models
3. **Documentation**: Keep anonymization logic well-documented
4. **Testing**: Implement comprehensive tests for PII detection
5. **Monitoring**: Set up alerts for mesh model failures
6. **Version Control**: Tag releases when exposing new models

## Troubleshooting

#### Common Issues

1. **Model not found in child**: Check access configuration and group assignment
2. **PII exposure**: Review anonymization logic and add tests
3. **Stale data**: Monitor upstream model runs in parent project
4. **Permission errors**: Verify Paradime project dependency configuration

#### Debug Commands

```bash
# Check available models from producer project
dbt list --resource-type model --output name --select customer_data_platform.*

# Validate model access and compilation
dbt compile --select +customer_data_platform.customers_anonymized

# Test mesh connectivity with specific producer models
dbt run --select customer_data_platform.customers_anonymized+

```

**Common Issues and Solutions:**

1. **"Model not found" errors**
   * Verify `dbt_loom.config.yml` configuration
   * Check that environment variables are set correctly
   * Ensure the Bolt schedule has run successfully in producer project
   * Confirm model has `access: public` in producer project
2. **API authentication errors**
   * Verify API credentials are correctly set at both workspace and user levels
   * Check API key permissions include "Bolt schedules metadata viewer"
   * Ensure API endpoint URL is correct
3. **Stale metadata**
   * Producer project must have successful Bolt schedule runs
   * Paradime fetches metadata from the specified schedule name
   * If producer models change, wait for next Bolt schedule run
4. **Model access denied**
   * Check model `access` configuration in producer project
   * Only `public` models are available through mesh
   * Verify model is in correct group with appropriate access level

This setup ensures that sensitive PII remains secure in the parent repository while providing rich, anonymized datasets for analytics in the child projects through Paradime's dbt mesh capabilities.

#### Related Docs:

* [Data Mesh Setup](/app-help/guides-new/data-mesh-setup)
* [Configure Project dependencies](/app-help/guides-new/data-mesh-setup/configure-project-dependencies)
* [Model groups](/app-help/guides-new/data-mesh-setup/model-groups)
* [Model access](/app-help/guides-new/data-mesh-setup/model-access)
* [Creating Schedules](/app-help/documentation/bolt/creating-schedules)
* [Environment Variables](/app-help/documentation/settings/environment-variables)


# Concepts

<table data-card-size="large" data-view="cards"><thead><tr><th></th><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th><th data-hidden data-card-cover data-type="files"></th></tr></thead><tbody><tr><td><p><strong>Working With Git</strong></p><p><br>Easy git for beginners and unlimited flexibility for power users.</p></td><td></td><td></td><td><a href="/pages/X9wVyyKKCt5e6UZyqi8h">/pages/X9wVyyKKCt5e6UZyqi8h</a></td><td><a href="/files/atkblDcI1WeK9D5G8Bma">/files/atkblDcI1WeK9D5G8Bma</a></td></tr><tr><td><strong>dbt™ Fundamentals</strong></td><td>Learn the dbt™ fundamentals with Paradime's tutorials and resources.</td><td></td><td><a href="/pages/nrpT8057QYoosgT0GA90">/pages/nrpT8057QYoosgT0GA90</a></td><td><a href="/files/NZVbG5GXkyi8jvM60abM">/files/NZVbG5GXkyi8jvM60abM</a></td></tr><tr><td><strong>How --defer works</strong></td><td>Learn how to continuously develop using production data and schema</td><td></td><td><a href="/pages/5AKNUFFsdhuLXZc4WTgI">/pages/5AKNUFFsdhuLXZc4WTgI</a></td><td><a href="/files/aDYxqPN0L17OxfsxWLcs">/files/aDYxqPN0L17OxfsxWLcs</a></td></tr><tr><td>Paradime Fundamentals</td><td>Learn the ins and outs of Paradime to perform analytics engineering with confidence.</td><td></td><td><a href="/pages/VutlficzORPuZtmHi2kB">/pages/VutlficzORPuZtmHi2kB</a></td><td><a href="/files/WCL21O2GM05oyXvwdRSm">/files/WCL21O2GM05oyXvwdRSm</a></td></tr></tbody></table>


# Working with Git

Paradime Git Integration: Manage dbt™ project version control with Git. Utilize comprehensive Git features for seamless collaboration.

Git integration in Paradime provides comprehensive version control for your dbt™ projects, ensuring safe collaboration and code management. Whether you're new to Git or an experienced user, Paradime offers the right interface for your workflow.

### Git Modes in Paradime

Paradime supports two Git interfaces designed for different user needs:

* [**Git Lite**](/app-help/concepts/working-with-git/git-lite): Simplified interface with built-in safeguards for Git beginners.
* [**Git Advanced**](/app-help/concepts/working-with-git/git-advanced): Full-featured Git interface for experienced users who need complete control

Choose your preferred mode in [IDE Preferences](/app-help/documentation/code-ide/user-interface/ide-preferences#git-experience-modes).

{% hint style="info" %}
**Git CLI Support**

The Git command line interface is available in the Paradime [Integrated Terminal](/app-help/documentation/code-ide/terminal) regardless of your selected Git mode.
{% endhint %}

***

### Git Workflow and & Management

<table data-view="cards"><thead><tr><th></th><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>🔒 Read-Only Branches</strong></td><td>Learn how Paradime protects your main branch and proper workflow practices.</td><td></td><td><a href="/pages/JCPTJEVC2iLYLTtZ7iru">/pages/JCPTJEVC2iLYLTtZ7iru</a></td></tr><tr><td><strong>🔀 Merge Conflicts</strong></td><td>Resolve conflicts with Paradime's visual resolution tools.</td><td></td><td><a href="/pages/3VOZmDVs0m2Zo6Nv5CmT">/pages/3VOZmDVs0m2Zo6Nv5CmT</a></td></tr><tr><td><strong>🔍 GitLens</strong></td><td>Enhanced visibility with blame annotations, history, and search capabilities.</td><td></td><td></td></tr><tr><td><strong>🗑️ Delete Branches</strong></td><td>Clean up merged branches to keep your workspace organized.</td><td></td><td><a href="/pages/wHcMsMMRgDOqgTizdfAm">/pages/wHcMsMMRgDOqgTizdfAm</a></td></tr></tbody></table>

***

### Security & Administration

<table data-view="cards"><thead><tr><th></th><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>✍️ Signed Commits</strong></td><td>Configure SSH key-based commit signing for enhanced security.</td><td></td><td><a href="/pages/4Gr8Qz1PuSmsMmG8Xttu">/pages/4Gr8Qz1PuSmsMmG8Xttu</a></td></tr><tr><td><strong>🛡️ GitHub Branch Protection</strong></td><td>Set up branch protection rules and enforce review processes.</td><td></td><td><a href="/pages/ahytUG2mJt3UoyC9XDG1">/pages/ahytUG2mJt3UoyC9XDG1</a></td></tr></tbody></table>


# Git Lite

Simplified Git interface for dbt™ projects with built-in safeguards. Perfect for users starting their Git journey while ensuring proper workflow practices.

Git Lite simplifies version control for dbt™ projects. It provides a purpose-built Git workflow that keeps you in sync with your main branch, perfect for users new to Git or those who prefer a streamlined interface.

{% hint style="info" %}
You can choose between Git Lite and Git Advanced in your [IDE preferences](/app-help/documentation/code-ide/user-interface/ide-preferences#git-workflow-preferences)
{% endhint %}

***

## Branch Management

### Create or switch Branches <a href="#create-or-switch-branches" id="create-or-switch-branches"></a>

Create a new branch to start working on features or fixes without affecting the main branch.

**Step-by-Step Instructions:**

1. **Click the Git icon** in the left panel of the Code IDE
2. **Select "Create Branch"** from the menu
3. **Enter your branch name** (use descriptive names like "feature-user-dashboard")
4. **Click "Create Branch"** to create and switch to your new branch

{% @arcade/embed url="<https://app.arcade.software/share/ZL4fNCjbrPyTR6LMT17R>" flowId="ZL4fNCjbrPyTR6LMT17R" %}

{% hint style="info" %}
Paradime will always create a new branch from your remote main/master branch. If switching to an existing branch Paradime will automatically switch branch using your remote selected branch.
{% endhint %}

### Switching Branches

Move between different branches to work on various features or review changes.

**Step-by-Step Instructions:**

1. **Click the Git Lite icon** (🪄) in the left panel of the Code IDE
2. **Click the branch dropdown** to see available branches
3. **Use the search bar** to find your branch if you have many
4. **Select the branch name** to switch to it

{% @arcade/embed url="<https://app.arcade.software/share/yw5rlZw2s7makCwUWdCr>" flowId="yw5rlZw2s7makCwUWdCr" %}

{% hint style="info" %}
Branch list updates automatically every 30 minutes. Click the refresh button to update manually.
{% endhint %}

***

### Working with Changes

### Keeping Branches Updated

Git Lite automatically checks if your branch is behind your remote branch and/or your main/master branch. This check happens:

* After each commit you make
* When you click the refresh button manually
* When you switch between branches

**Step-by-Step Instructions:**

1. **Click the Git icon** in the left panel of the Code IDE
2. Click the **refresh** icon in the status section

{% @arcade/embed url="<https://app.arcade.software/share/tkKbvWfgRk6cO3Adbxgd>" flowId="tkKbvWfgRk6cO3Adbxgd" %}

### Committing Changes

When you have file changes you want to commit, follow these instructions.

**Step-by-Step Instructions:**

1. **Click the Git icon** in the left panel of the Code IDE
2. **Review your changed files** shown in the panel
3. **Enter a clear commit message** describing what you changed
4. **Click "Commit and Push"** to save and upload your changes

{% hint style="info" %}
You can also discard all changes by clicking on the **`Discard Changes`** button.
{% endhint %}

{% @arcade/embed url="<https://app.arcade.software/share/AZDaPJ6I1gZv3Pn3xG4F>" flowId="AZDaPJ6I1gZv3Pn3xG4F" %}

{% hint style="info" %}
**How it works**

**Git Lite automatically:**

* Commits your changes with the provided message
* Pushes changes to your remote branch
* Pulls and merges remote changes if your branch is behind
* Keeps you in sync without manual intervention
  {% endhint %}

### Open Pull Request <a href="#open-pull-request" id="open-pull-request"></a>

Create a pull request to propose merging your changes into the main branch.

**Step-by-Step Instructions:**

1. **Click the Git icon** in the left panel of the Code IDE
2. **Ensure all changes are committed** by clicking "Commit and Push" first
3. **Click "Open Pull Request"** (only enabled when branch is clean)
4. **Complete the pull request** in your Git provider interface

<figure><img src="/files/i1hIeu31gBBwF0VDQbi7" alt="" width="237"><figcaption></figcaption></figure>

### Reverting Commits

Undo your last commit if you need to fix something or made an error.

**Step-by-Step Instructions:**

1. **Click the Git icon** in the left panel of the Code IDE
2. **Click "Revert Commit"** in the top right corner of the Git Lite panel
3. **Review the commit details** shown in the dialog
4. **Click "Continue"** to revert the last commit

You will see in the dialog the commit details include the commit message, click on `continue` to complete the operation.

<figure><img src="/files/DlLZensN3rAvDoC1kbfy" alt="" width="235"><figcaption></figcaption></figure>

### Merge Conflicts <a href="#merge-conflicts" id="merge-conflicts"></a>

When merge conflicts occur, Git Lite provides tools to help you resolve them easily.

{% hint style="info" %}
For detailed instructions on resolving merge conflicts, see our [merge conflict](/app-help/concepts/working-with-git/merge-conflicts) documentation.
{% endhint %}


# Git Advanced

Paradime Git Advanced: Advanced Git features for complex dbt™ projects. Enhance code management and collaboration.

Git Advanced provides a full-featured Git interface for experienced users who need complete control over their version control workflow. It offers all standard Git operations including selective staging, advanced branching, and detailed commit management.

{% hint style="info" %}
Choose Git Advanced if you're comfortable with Git concepts, need granular control over staging files, or want to perform complex Git operations. You can switch from Git Lite anytime in [IDE Preferences](/app-help/documentation/code-ide/user-interface/ide-preferences).
{% endhint %}

## Create or switch Branches <a href="#create-or-switch-branches" id="create-or-switch-branches"></a>

To create a new branch click on the branch name on the bottom-left of your screen, from the menu select `Create a new branch` and add the new branch name. Press `Enter` to create your new branch.

{% hint style="info" %}
Make sure to Publish to remote the new branch you just created.
{% endhint %}

{% @arcade/embed url="<https://app.arcade.software/share/uWAtf8lqRU4EOhFZDEPG>" flowId="uWAtf8lqRU4EOhFZDEPG" %}

{% hint style="warning" %}
The default Git behavior is to create a new branch form the currently checked out branch. If you want to create a new branch based on the latests version of Main/Master, make sure to:

1. First select the Main/Master branch
2. Update your local Main/Master by doing a Git Pull
3. Now create the branch from the UI
   {% endhint %}

## Commit and Push your Changes <a href="#commit-and-push-your-changes" id="commit-and-push-your-changes"></a>

### Staging changes <a href="#staging" id="staging"></a>

When using Git Advance before committing you changes you need to select which files you want to stage first. This is the equivalent of the Git command [`git add`](https://www.atlassian.com/git/tutorials/saving-changes). Here you can choose to stage all you changed files or selectively choose the file to be included in the next commit.

### Commit your changes and push to remote

After you staged you changes you can now enter a commit message and click on the commit button on the top of the Git Advance panel option or from the More options context menu.

{% @arcade/embed url="<https://app.arcade.software/share/13PjzeH0kpzcjltO09YR>" flowId="13PjzeH0kpzcjltO09YR" %}

## Open Pull Request <a href="#open-pull-request" id="open-pull-request"></a>

Now that you have created your commits and pushed your changes you can use the Open Pull Request option to go directly to your Git Provider interface and open a PR for review.

{% @arcade/embed url="<https://app.arcade.software/share/G5Zn3IjRVyqZz1guJskH>" flowId="G5Zn3IjRVyqZz1guJskH" %}

## Merge Conflicts <a href="#merge-conflicts" id="merge-conflicts"></a>

In the event where you have merge conflicts, Paradime can help you resolve them. See our merge conflict documentation for details.

{% content-ref url="/pages/3VOZmDVs0m2Zo6Nv5CmT" %}
[Merge Conflicts](/app-help/concepts/working-with-git/merge-conflicts)
{% endcontent-ref %}


# Read Only Branches

Paradime Read-Only Branches: Secure your dbt™ projects with read-only branches. Ensure integrity and control over code changes.

{% hint style="warning" %}
While developing your dbt™️ project, changes/edits should never made against your remote default branch (usually called main or master).

When working with git a new branch should be created in order to make any changes to your project so that a Pull Request can be opened to merge your changes into your default branch.
{% endhint %}

To prevent users to commit directly again main or master, whether you are using [Git Lite](/app-help/documentation/code-ide/left-panel/git-lite) or [Git Advanced](/app-help/concepts/working-with-git/git-advanced) the default branch will be always protected in Paradime meaning that you wont be able to make edits to files in your dbt™️ project.

If you see the prompt below when trying to edit a file, simply create a new branch or switch to an existing one.

<figure><img src="/files/LAaNPLWuCn7f8F7ZEMJ0" alt=""><figcaption></figcaption></figure>


# Merge Conflicts

Paradime Merge Conflicts: Resolve merge conflicts in dbt™ projects effortlessly. Ensure smooth collaboration and code integration.

Merge conflicts arise when two branches have changes that Git cannot automatically merge. This situation commonly occurs when:

* Two branches have changes in the same part of a file.
* One branch has changes, and the other branch deletes the same file.
* Both branches have changes, and Git can't determine which changes should take precedence.

### **Using Merge Conflicts in Paradime**

To use merge conflict, follow the step below:

#### **1. Open a Pull Request**

Click the git icon on the right panel to open the Git User Interface, and then commit and push your changes.\
\
Note: The visual tutorial below is for [GitLite](/app-help/documentation/code-ide/left-panel/git-lite) users, but the merge conflict feature works for[ Git Advanced](/app-help/concepts/working-with-git/git-advanced) users as well.

{% @arcade/embed url="<https://app.arcade.software/share/j9gbtzjCmzoeh7IN09ND>" flowId="j9gbtzjCmzoeh7IN09ND" %}

#### **2. Resolve Merge Conflict(s)**

Paradime offers various options to resolve merge conflicts:

* **Abort Merge -** Cancel the entire merge process if conflicts are significant or challenging to resolve.
* **Accept Current Changes -** Accept and apply the changes made in your branch.
* **Accept Incoming Changes -** Accept and apply the incoming changes from the target branch.
* **Accept Both Changes -** Combine and apply both sets of changes – yours and the incoming changes from the target branch.
* **Compare Changes -** Visually compare the changes made in your branch with incoming changes from the target branch.

{% @arcade/embed url="<https://app.arcade.software/share/MjA8Co7lz6yIcIdhTXE7>" flowId="MjA8Co7lz6yIcIdhTXE7" %}

#### **3. Commit and Resolve**

Once you've decided how to resolve the merge conflict, proceed with your pull request by clicking the git icon in the right panel, "Commit and Resolve" and then "Open Pull Request".

{% @arcade/embed url="<https://app.arcade.software/share/As7tOqpbdyqgbABNPMGm>" flowId="As7tOqpbdyqgbABNPMGm" %}

‍


# Delete Branches

Clean up your dbt™ project repository by safely deleting merged branches using one-click pruning or manual deletion to keep your workspace organized

Keep your dbt™ project repository organized by deleting branches that are no longer needed. Paradime offers two approaches: [automated pruning](#one-click-branch-pruning) that removes multiple merged branches at once, and [manual deletion](#manual-branch-deletion) for removing specific branches individually.

***

### One-Click Branch Pruning

One-Click Branch Pruning automatically identifies and removes local branches that have been merged and deleted on remote. This feature eliminates the need for manual cleanup and helps maintain an organized workspace.

{% hint style="success" %}
**Benefits**

* **Automatic Detection**: Instantly identifies branches that were deleted on remote
* **Safe Cleanup**: Only removes branches that have been properly merged
  {% endhint %}

Branch pruning is available for both [GitLite](#git-lite) and [Git Advanced](#git-advanced) users. See below step by step instructions:

{% tabs %}
{% tab title="Git Lite" %}
**Step-by-Step Instructions:**

1. **Open Git Lite panel** in the left sidebar
2. **Click the scissor icon** to access branch pruning
3. **Review the branches** that will be removed (shows exactly which branches will be deleted)
4. **Confirm the cleanup** to remove all identified branches

{% @arcade/embed url="<https://app.arcade.software/share/V070TnrZOPJjTDqVc24f>" flowId="V070TnrZOPJjTDqVc24f" %}
{% endtab %}

{% tab title="Git Advanced" %}
**Step-by-Step Instructions:**

1. **Open Git panel** in the left sidebar
2. **Click the three-dot menu (⋯)** on top left of side bar
3. **Select "Prune deleted branches" from the dropdown**
4. **Review the branches** that will be removed (shows exactly which branches will be deleted)
5. **Confirm the cleanup** to remove all identified branches

{% @arcade/embed url="<https://app.arcade.software/share/7MCGhrnJKTbBBEX3ztec>" flowId="7MCGhrnJKTbBBEX3ztec" %}
{% endtab %}
{% endtabs %}

{% hint style="info" %}
**Smart Detection**

The pruning feature shows you exactly which branches will be removed before deletion, eliminating guesswork and preventing accidental removal of important branches.
{% endhint %}

***

### Manual Branch Deletion

For Git Advanced users who need to delete specific branches individually or remove branches that haven't been merged yet.

**Step-by-Step Instructions:**

1. **Open the Git Advanced panel** in the left sidebar
2. **Click the three-dot menu (⋯)** on top left of side bar
3. **Select "Delete Branches"** from the dropdown
4. **Choose the specific branch** you want to delete from the list
5. **Confirm the deletion**

{% @arcade/embed url="<https://app.arcade.software/share/iqZsKz3tAuDg5oyfZo9a>" flowId="iqZsKz3tAuDg5oyfZo9a" %}

{% hint style="danger" %}
A branch can only be deleted from the UI if it has already been pushed and merged with the remote branch. To force delete an unmerged branch, use the terminal command: **`git branch -D <branch_name>`**
{% endhint %}


# GitLens

Enhanced Git visibility and history tracking in Paradime with GitLens integration. Explore commit history, blame annotations, and file changes with powerful search capabilities.

GitLens enhances your Git experience in Paradime by providing powerful visualization and navigation tools for understanding code history, authorship, and changes. This integration brings advanced Git insights directly into your dbt™ development workflow.

GitLens is available through a dedicated icon in the left sidebar of the Code IDE and includes three main functionalities that are commonly referred to as "Git Blame" or "blame annotations":

* [File History](#file-history) for tracking how entire files have evolved over time
* [Line History](#line-history) for seeing who changed specific lines of code
* [Search & Compare](#search-and-compare) for finding specific commits across your project

{% hint style="info" %}
GitLens features are available in both [Git Lite](/app-help/concepts/working-with-git/git-lite) and [Git Advanced](/app-help/concepts/working-with-git/git-advanced) modes, providing the same functionality regardless of your chosen Git interface.
{% endhint %}

***

### File History

File History shows a chronological list of all commits that affected a selected file, helping you understand how entire files have evolved over time. This feature is useful for seeing the overall development pattern of specific files and understanding major changes made by your team.

**Step-by-Step Instructions:**

1. **Select a file** in the Code IDE
2. **Click the GitLens icon** in the left sidebar of the Code IDE
3. **View "File History" tab** in the GitLens panel to see chronological list of all commits
4. **Click on specific commits** to view detailed changes made in that commit

{% @arcade/embed url="<https://app.arcade.software/share/3JonU6LhgEztHWgzjFkG>" flowId="3JonU6LhgEztHWgzjFkG" %}

### Line History

Line History shows who made changes to specific lines of code and when they were made through inline annotations that appear next to each line. This feature helps you understand the context behind code changes and identify the author of specific modifications for collaboration or debugging purposes.

**Step-by-Step Instructions:**

1. **Open any file** in the Code IDE
2. **Click a line of code** to see inline annotations showing the most recent author and commit date
3. **Hover over annotations** to see additional details about the most recent commit
4. **For full line history** (beyond just the most recent commit):
   * **Click the GitLens icon** in the left sidebar of the Code IDE
   * **Reselect the line** to view complete change history for that specific line

{% @arcade/embed url="<https://app.arcade.software/share/9YiqmuOAn2O6Z3grzJaL>" flowId="9YiqmuOAn2O6Z3grzJaL" %}

### Search & Compare

Search & Compare allows you to quickly find commits, changes, or contributions across your entire project, as well as compare different branches or references. This feature eliminates the need to manually browse through commit history when looking for specific changes or contributions.

**Step-by-Step Instructions:**

1. **Click the GitLens icon** in the left sidebar of the Code IDE
2. **Navigate to "Search & Compare" tab** in the GitLens panel
3. **Choose between two options**:

**Option 1: Search Commits**

* **Select your search type** (Author, Message, or File)
* **Enter search terms** (e.g., author name "fabio" or message keyword "fix")
* **Review search results** and click on commits to view detailed information

**Option 2: Compare Reference**

* **Select the first reference** (branch, tag, or ref) you want to compare
* **Select the second reference** (branch, tag, or ref) to compare against
* **Review the comparison** to see differences between the two references

{% @arcade/embed url="<https://app.arcade.software/share/utUevOUSrxsn5Mxx3gfk>" flowId="utUevOUSrxsn5Mxx3gfk" %}


# Configuring Signed Commits on Paradime with SSH Keys

This guide explains how to set up SSH key-based commit signing on Paradime, which enhances the security and verification of your Git commits.

This guide explains how to set up SSH key-based commit signing on Paradime, which enhances the security and verification of your Git commits.

### Why Sign Your Commits?

Signing your commits verifies that you are the author of your code changes and helps maintain the integrity of your codebase by preventing commit spoofing.

{% hint style="info" %}

#### Prerequisites

* Paradime IDE access
* Git repository initialized in your Paradime workspace
* GitHub account (for adding your signing key)
  {% endhint %}

### Setup Instructions

#### Step 1: Create the Setup Script

In your Paradime IDE, create a new file called `setup_git_signed_commits.sh` with the following content:

```bash
#!/bin/bash

# Function to check if we're in a git repository
check_git_repo() {
    if ! git rev-parse --git-dir > /dev/null 2>&1; then
        echo "Error: Not a git repository"
        exit 1
    fi
}

# Function to generate SSH key
generate_ssh_key() {
    local key_comment=$1
    
    if [ -f ~/.ssh/git_signing_key ]; then
        echo "Warning: SSH key git_signing_key already exists"
        read -p "Do you want to overwrite it? (y/n) " -n 1 -r
        echo
        if [[ ! $REPLY =~ ^[Yy]$ ]]; then
            echo "Aborting..."
            exit 1
        fi
    fi
    
    ssh-keygen -t ed25519 -C "$key_comment" -f ~/.ssh/git_signing_key -N ""
    
    # Set correct permissions
    chmod 600 ~/.ssh/git_signing_key
    chmod 644 ~/.ssh/git_signing_key.pub
}

# Function to configure git
configure_git() {
    echo -e "\nSetting local git configuration to use the generated signing key.."
    
    git config gpg.format ssh
    git config user.signingkey "~/.ssh/git_signing_key.pub"
    git config commit.gpgsign true
    
    echo -e "Git configuration complete!\n\n"
}

# Function to display public key
display_key() {
    echo "Here's your public key to add to GitHub:"
    echo "----------------------------------------"
    cat ~/.ssh/git_signing_key.pub
    echo "----------------------------------------"
    echo "Add this key to GitHub by visiting: https://github.com/settings/keys"
    echo "Make sure to choose the key type as 'Signing Key' when adding it. Once done, your setup is complete."
}

# Main script
main() {
    check_git_repo
    
    # Get user input
    read -p "Enter a comment for your key (e.g., your name, email, etc): " key_comment
    
    # Setup steps
    generate_ssh_key "$key_comment"
    configure_git
    display_key
}

# Run main function
main
```

#### Step 2: Make the Script Executable

Open your Paradime terminal and run the following command to make the script executable:

```bash
chmod +x setup_git_signed_commits.sh
```

#### Step 3: Run the Setup Script

Execute the script by running:

```bash
./setup_git_signed_commits.sh
```

When prompted, enter a comment for your key (typically your name and email address).

#### Step 4: Add Your Signing Key to GitHub

1. Copy the public key that is displayed in the terminal output
2. Go to your GitHub account settings: <https://github.com/settings/keys>
3. Click "New SSH key"
4. Choose "Signing Key" as the key type
5. Paste your public key in the provided field
6. Give your key a descriptive title
7. Click "Add SSH key"

#### Step 5: Verification

Your setup is now complete! Every new commit you make in this repository will be automatically signed with your SSH key.

You can verify a signed commit by viewing it on GitHub, where you should see a "Verified" badge next to properly signed commits.


# GitHub Branch Protection Guide: Preventing Direct Commits to Main

Learn to secure your codebase with GitHub branch protection rules to prevent direct pushes to main, enforce pull request reviews.

## Introduction

Branch protection rules are essential for maintaining code quality and preventing accidental or unauthorized changes to important branches like `main`. This guide will walk you through setting up branch protection rules in GitHub and ensuring they're properly enforced across your organization.

### Setting Up Branch Protection Rules

#### Basic Branch Protection

1. Navigate to your repository on GitHub
2. Click on "Settings" in the top navigation bar
3. In the left sidebar, click on "Branches"
4. Under "Branch protection rules," click "Add rule"
5. In the "Branch name pattern" field, enter `main` (or your default branch name)
6. Check the following options:
   * "Require a pull request before merging"
   * "Require approvals" (set the number of required reviewers, typically at least 1)
   * "Dismiss stale pull request approvals when new commits are pushed"
   * "Require status checks to pass before merging"
   * "Require branches to be up to date before merging"

<figure><img src="/files/QhdTiBeL84Q64rH9t8pA" alt=""><figcaption></figcaption></figure>

Under "Rules applied to everyone including administrators", check "**Do not allow bypassing the above settings**"

<figure><img src="/files/cwbVtZvir0eN5JPRTVHr" alt=""><figcaption></figcaption></figure>

#### Advanced Protection Settings

For stronger protection:

1. Enable "Include administrators" to apply rules to everyone
2. Check "Restrict who can push to matching branches" if you want only specific teams/people to merge PRs
3. Enable "Allow force pushes" only for specific people/teams if absolutely necessary

### Enforcing Organization-Wide Branch Protection

To ensure consistent protection across all repositories:

#### Using Organization Repository Rules

1. Navigate to your GitHub organization
2. Click on "Settings" in the top navigation menu
3. In the left sidebar, click on "Repository rules"
4. Click "New rule"
5. Name your rule (e.g., "Main Branch Protection")
6. Under "Branch protections", configure the same settings as above
7. Set the rule to apply to:
   * All repositories, or
   * Repositories matching specific criteria (e.g., visibility, topics)
8. Click "Create rule"

#### Using GitHub Enterprise Policies (For Enterprise Accounts)

If you have GitHub Enterprise:

1. Go to your enterprise account settings
2. Navigate to "Policies" > "Repository"
3. Under "Repository policies", scroll to "Branch protection rules"
4. Enable "Require branch protection rules" and configure the default settings
5. Save your changes

### Verifying Branch Protection

To ensure your protections are working correctly:

1. Try pushing directly to the main branch from a local repository

   ```bash
   git checkout maingit commit -m "Test commit"git push
   ```

   This should be rejected with an error message

<figure><img src="/files/QA8PKXFQMbh5WF1kuZpy" alt=""><figcaption></figcaption></figure>

1. Create a new branch, commit changes, and open a pull request

   ```bash
   git checkout -b feature-branchgit commit -m "Test PR"
   git push -u origin feature-branch
   ```

   Then create a PR in the GitHub UI
2. Attempt to merge the PR without meeting requirements (this should be blocked)

{% hint style="warning" %}
**Troubleshooting Common Issues**

* **Settings not applying**: Verify "Include administrators" is checked
* **Bypassed protections**: Check that "Do not allow bypassing the above settings" is enabled
* **Repository-specific exceptions**: Review organization rules for conflicts
* **Branch deletion issues**: Enable "Restrict deletions" in branch protection settings
  {% endhint %}

### Best Practices

* Protect all production branches (`main`, `production`, etc.)
* Require at least one review for all PRs
* Configure required status checks for CI/CD pipelines
* Consider requiring signed commits for additional security
* Regularly audit branch protection settings across repositories
* Document your branch protection strategy for team reference

By implementing these protections, you'll help ensure code quality and prevent accidental deployments to critical branches.


# GitHub Merge Queue

Automate GitHub pull request merging with merge queue CI/CD integration. Prevent code conflicts and failed builds using automated testing workflows.

## Overview

GitHub Merge Queue is a feature that helps maintain code quality and prevent integration issues by automatically testing pull requests together before merging them into the main branch. This documentation covers how to set up and configure merge queue with CI/CD workflows.

### What is Merge Queue?

Merge Queue ensures that:

* Pull requests are tested in the exact state they'll be in after merging
* Multiple PRs can be batched and tested together
* Failed tests prevent problematic code from reaching the main branch
* Integration issues are caught before merge, not after

### Prerequisites

Before setting up merge queue, ensure you have:

* Repository admin permissions
* [A Paradime Bolt Turbo CI schedule configured](/app-help/documentation/bolt/creating-schedules/templates/ci-cd-templates/test-code-changes-on-pull-requests)
* [Paradime API credentials](/app-help/developers/generate-api-keys-legacy#generate-a-new-set-of-api-keys) set in your GitHub repo
* A CI/CD workflow already configured
* Required status checks defined for your repository

## Setup Process

### Step 1: Create a Merge Queue Workflow

Create a new GitHub Actions workflow file (e.g., `.github/workflows/merge-queue-ci.yml`) that triggers on `merge_group` events:

```yaml
name: Paradime Turbo CI
on:
  merge_group:

concurrency:
  group: ${{ github.workflow }}-${{ github.ref }}
  cancel-in-progress: true

jobs:
  paradime-turbo-ci:
    name: Paradime Turbo CI
    runs-on: ubuntu-latest

    steps:
      - name: Checkout repository
        uses: actions/checkout@v3

      - name: Set up Python
        uses: actions/setup-python@v5
        with:
          python-version: '3.12'

      - name: Install paradime-io
        run: |
          python -m pip install --upgrade pip
          pip install paradime-io==4.6.0

      - name: Run Paradime CLI command
        env:
          PARADIME_API_ENDPOINT: ${{ secrets.PARADIME_API_ENDPOINT }}
          PARADIME_API_KEY: ${{ secrets.PARADIME_API_KEY }}
          PARADIME_API_SECRET: ${{ secrets.PARADIME_API_SECRET }}
          SCHEDULE_NAME: >-
            <YOUR_TURBO_CI_SCHEDULE_NAME>
        run: paradime bolt run "${{ env.SCHEDULE_NAME }}" --branch "${{ github.sha }}" --wait
```

#### **Key Configuration Points:**

* **Trigger**: `merge_group` event ensures the workflow runs when PRs are added to the merge queue
* **Concurrency**: Prevents multiple instances of the same workflow from running simultaneously
* **Schedule Name**: Replace `YOUR_TURBO_CI_SCHEDULE_NAME` with your actual Turbo CI schedule name
* **Secrets**: Ensure all required API credentials are configured in repository secrets

### Step 2: Configure Repository Settings

Navigate to your repository's Settings → General → Pull Requests:

1. **Enable Merge Queue**:
   * Check "Require merge queue"
   * Configure merge queue settings according to your needs:
     * **Merge method**: Choose squash, merge commit, or rebase
     * **Build concurrency**: Set maximum number of PRs to test simultaneously
     * **Merge queue grouping**: Configure batching behavior
2. **Configure Required Status Checks**:
   * Go to Settings → Branches
   * Edit your branch protection rule for the main branch
   * Add "Paradime Turbo CI" as a required status check
   * ⚠️ **Important**: Set the source to "any source" since the status check comes from both Paradime and the GitHub Actions workflow

<figure><img src="/files/vmBrGPmMAnPhskvCMbLo" alt=""><figcaption></figcaption></figure>

## Workflow Behavior

### How It Works

1. **PR Creation**: Developer creates a pull request
2. **Queue Addition**: When ready to merge, PR is added to the merge queue
3. **Batch Testing**: GitHub creates a temporary merge commit combining the PR with the target branch
4. **CI Execution**: The merge queue workflow runs against this temporary commit
5. **Merge Decision**:
   * ✅ If tests pass: PR is automatically merged
   * ❌ If tests fail: PR is removed from queue and developer is notified

### Queue Management

* **Automatic Batching**: Multiple PRs can be tested together for efficiency
* **Failure Handling**: Failed PRs are automatically removed without affecting others
* **Priority**: PRs are processed in the order they were added to the queue

## Best Practices

### Workflow Design

* **Keep Tests Fast**: Merge queue workflows should be optimized for speed
* **Use Caching**: Implement dependency caching to reduce build times
* **Parallel Execution**: Structure jobs to run in parallel where possible
* **Clear Naming**: Use descriptive job and step names for easier debugging

### Repository Configuration

* **Require Reviews**: Combine merge queue with required code reviews
* **Branch Protection**: Use comprehensive branch protection rules
* **Status Checks**: Only require essential checks to avoid bottlenecks
* **Documentation**: Keep team documentation updated on merge queue usage

#### Team Workflow

* **Clear Guidelines**: Establish when to use merge queue vs. direct merge
* **Monitor Queue**: Regularly check queue status and resolve issues promptly
* **Communication**: Notify team members when queue is blocked

## Troubleshooting

### Common Issues

**Workflow Not Triggering**

* Verify `merge_group` trigger is correctly configured
* Check that the workflow file is in the correct location
* Ensure branch protection rules include the workflow as a required check

**Status Check Failures**

* Confirm all required secrets are configured
* Verify schedule name matches your Turbo CI configuration
* Check API endpoint and credentials are correct

**Queue Bottlenecks**

* Review concurrency settings
* Optimize workflow performance
* Consider adjusting batch size settings

#### Debugging Tips

* Check the "Actions" tab for detailed workflow logs
* Review merge queue status in the PR interface
* Monitor repository insights for queue performance metrics

{% hint style="info" %}
GitHub Merge Queue provides a robust solution for maintaining code quality while enabling efficient team collaboration. By following this setup guide and best practices, your team can implement a reliable merge process that scales with your development workflow.\
\
For additional support or advanced configurations, consult the [GitHub Documentation](https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/configuring-pull-request-merges/managing-a-merge-queue) or reach out to your DevOps team.
{% endhint %}


# dbt™ fundamentals

Learn the dbt basics with Paradime's clear, step-by-step guidance, making it straightforward to master dbt and start building with confidence.


# Getting started with dbt™

An introduction to dbt™ for beginners, covering core concepts, project structure,   models, sources, and testing to help you build reliable data transformations.

## Getting Started with dbt™

dbt™ (Data Build Tool) is a transformation framework that enables data teams to apply software engineering best practices to analytics workflows. This section provides a high-level introduction to dbt™, how it fits into the modern data stack, and how to structure your dbt™ projects effectively.

{% hint style="info" %}
If you're new to dbt™, this guide will walk you through the key concepts of dbt™
{% endhint %}

### Explore the Basics of dbt™

<table data-view="cards"><thead><tr><th></th><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th><th data-hidden data-card-cover data-type="files"></th></tr></thead><tbody><tr><td>📌 <strong>1. Introduction</strong></td><td>Understand how dbt™ improves data transformations, version control, and testing in the modern data stack.</td><td></td><td><a href="/pages/ZyVnVxEktyuFOweRC3s8">/pages/ZyVnVxEktyuFOweRC3s8</a></td><td></td></tr><tr><td>📁 <strong>2. Project Structure</strong></td><td>Learn how dbt™ projects are structured, including key files like <code>dbt_project.yml</code>, models, sources, and dependencies.</td><td></td><td><a href="/pages/iqLm8pH7yM1KGG5GYMXj">/pages/iqLm8pH7yM1KGG5GYMXj</a></td><td></td></tr><tr><td>🔗 <strong>4. Working with Sources</strong></td><td>Define and manage raw data sources using <code>sources.yml</code>, ensuring consistency, documentation, and data freshness.</td><td></td><td><a href="/pages/CGS8vexRhFuFkWIGCBec">/pages/CGS8vexRhFuFkWIGCBec</a></td><td></td></tr><tr><td>📊 <strong>3. Models and Transformations</strong></td><td>Discover how dbt™ models work, including ref(), materializations, and Jinja for dynamic transformations.</td><td></td><td><a href="/pages/ABZ1oYjwFeihn5VriirB">/pages/ABZ1oYjwFeihn5VriirB</a></td><td></td></tr><tr><td>🛠️ <strong>5. Testing Data Quality</strong></td><td>Explore dbt™ testing, including built-in and custom tests to validate data integrity and enforce business rules.</td><td></td><td><a href="/pages/qDHsH1hPCR8rcK8L31mm">/pages/qDHsH1hPCR8rcK8L31mm</a></td><td></td></tr></tbody></table>

{% hint style="info" %}
Prefer hands-on learning? Check out our [**Paradime 101 Guide**](/app-help/guides/paradime-101) for a step-by-step, interactive way to learn dbt™ and analytics engineering best practices—all for free.
{% endhint %}


# Introduction

Learn how dbt™ transforms SQL-based analytics by centralizing data transformations in the warehouse. This guide covers dbt’s role in the ELT stack, its benefits, and how it improves data reliability,

dbt (data build tool) transforms how data teams work by bringing software engineering practices to data transformation. It enables analysts and engineers to build reliable, modular, and tested data pipelines using simple SQL.

### What is dbt?

dbt is a transformation framework that works with your existing data warehouse. You write SQL SELECT statements, and dbt handles the complexity of turning them into tables and views while managing dependencies between models.

```sql
-- A simple dbt model example
SELECT 
    orders.id as order_id,
    orders.status,
    customers.name as customer_name,
    customers.email
FROM raw_data.orders
JOIN raw_data.customers ON orders.customer_id = customers.id
WHERE orders.status != 'cancelled'
```

This SQL becomes a fully-managed transformation with version control, testing, and documentation.

***

### Why dbt Matters

Modern data teams face growing complexity:

| Challenge                 | Impact                                       | dbt Solution                                 |
| ------------------------- | -------------------------------------------- | -------------------------------------------- |
| Scattered transformations | Inconsistent business logic, duplicated work | Centralized repository of transformations    |
| Manual SQL processes      | Errors, slow iterations, impossible to audit | Automated, repeatable transformation runs    |
| No testing                | Data quality issues affecting decisions      | Built-in testing framework                   |
| Poor documentation        | Knowledge silos, difficult onboarding        | Auto-generated, always-current documentation |

***

### The ELT Approach: Why It Matters

Modern data workflows have shifted from ETL (Extract, Transform, Load) to ELT (Extract, Load, Transform), and dbt is specifically built for the "T" in ELT.

In the ELT approach:

1. Raw data lands in your warehouse without transformation
2. dbt runs SQL transformations directly in the warehouse
3. Analysts work with cleaned, tested data models

This approach leverages your warehouse's processing power and keeps all transformations in a single, manageable location.

***

### How dbt Works in Practice

Let's follow the journey of a typical dbt workflow:

1. **Define a Source**: Tell dbt where your raw data lives

   ```yaml
   # sources.yml
   sources:
     - name: raw_data
       tables:
         - name: orders
         - name: customers
   ```
2. **Create a Model**: Write SQL that transforms this data

   ```sql
   -- models/staging/stg_orders.sql
   SELECT
     id as order_id,
     customer_id,
     status,
     created_at,
     -- Additional business logic here
     CASE 
       WHEN status = 'shipped' THEN 'completed'
       WHEN status = 'processing' THEN 'in_progress'
       ELSE 'other'
     END as order_status_normalized
   FROM {{ source('raw_data', 'orders') }}
   ```
3. **Add Tests**: Ensure data quality

   ```yaml
   # schema.yml
   models:
     - name: stg_orders
       columns:
         - name: order_id
           tests:
             - unique
             - not_null
   ```
4. **Run dbt**: Transform and test your data

   ```bash
   dbt run
   dbt test
   ```
5. **Document & Share**: Automatically generate documentation

   ```bash
   dbt docs generate
   dbt docs serve
   ```

***

### The dbt Ecosystem

dbt fits into a modern data stack alongside other specialized tools:

* **Extraction tools** (Fivetran, Airbyte) bring data to your warehouse
* **dbt** transforms this data into analytics-ready models
* **BI tools** (Tableau, Looker) visualize the transformed data

This modular approach allows each tool to focus on what it does best, creating a more maintainable data platform.

***

### Key Benefits

* **Write just SQL**: No new language to learn
* **Version-controlled transformations**: Track changes with Git
* **Automated testing**: Ensure data quality
* **Self-documenting models**: Always up-to-date documentation
* **Development workflows**: Build and test locally before deploying
* **Modular design**: Reusable patterns and dependencies

By centralizing transformations in dbt, data teams can build more reliable data pipelines, collaborate more effectively, and spend more time on analysis instead of maintenance.

{% hint style="info" %}
**Want to start using dbt in Paradime for free?**

Check out our [Paradime 101 guide](/app-help/guides/paradime-101) to set up your first dbt project!
{% endhint %}


# Project Strucuture

Understand how dbt™ projects are structured for scalability and maintainability. This guide covers key components like models, sources, configuration files, and best practices for organizing SQL trans

A dbt project is a collection of files that define how raw data should be transformed into analytics-ready datasets. Understanding the structure of a dbt project helps you organize transformations effectively and collaborate with your team.

### Anatomy of a dbt Project

When you initialize a new dbt project, you'll see a directory structure like this:

```
dbt_project/
├── models/          # SQL transformations (core of your project)
├── analyses/        # One-off analytical queries
├── tests/           # Custom data tests
├── macros/          # Reusable SQL code blocks
├── snapshots/       # Historical data tracking definitions
├── seeds/           # CSV files to be loaded into the database
├── dbt_project.yml  # Project configuration
├── packages.yml     # External dependency definitions
└── README.md        # Project documentation
```

***

### Core Components

#### Models Directory

The `models/` directory contains SQL files that define your transformations. Each SQL file typically becomes a table or view in your data warehouse.

```sql
-- models/marts/customers.sql
SELECT
    c.customer_id,
    c.name,
    c.email,
    COUNT(o.order_id) as number_of_orders,
    SUM(o.amount) as total_order_value
FROM {{ ref('stg_customers') }} c
LEFT JOIN {{ ref('stg_orders') }} o ON c.customer_id = o.customer_id
GROUP BY 1, 2, 3
```

Models are often organized into subdirectories by function or data domain:

```
models/
├── staging/          # Initial cleaning/renaming of source tables
├── intermediate/     # Intermediate transformations
└── marts/            # Business-level presentation models
```

#### Configuration Files

Configuration files define project-wide settings and metadata:

**dbt\_project.yml**

This is the central configuration file for your dbt project. It defines:

* Project name and version
* Profile to use for database connections
* Model materialization settings
* Directory configurations

```yaml
name: 'ecommerce'
version: '1.0.0'
config-version: 2

profile: 'ecommerce'

model-paths: ["models"]
seed-paths: ["seeds"]
test-paths: ["tests"]
macro-paths: ["macros"]
snapshot-paths: ["snapshots"]

models:
  ecommerce:
    staging:
      +materialized: view
    intermediate:
      +materialized: view
    marts:
      +materialized: table
```

**packages.yml**

This file defines external dbt packages that your project depends on:

```yaml
yamlCopypackages:
  - package: dbt-labs/dbt_utils
    version: 0.8.0
  - package: calogica/dbt_expectations
    version: 0.5.0
```

#### Source Definitions

Sources represent the raw data tables in your warehouse. They're defined in YAML files (typically named `sources.yml`) within the models directory:

```yaml
version: 2

sources:
  - name: raw_data
    schema: raw
    tables:
      - name: customers
        columns:
          - name: id
            tests:
              - unique
              - not_null
      - name: orders
        columns:
          - name: id
            tests:
              - unique
              - not_null
```

***

### Model Organization Patterns

There's no single "right way" to organize your dbt project, but here are common patterns that work well:

#### Layered Approach

This approach organizes models by their purpose in the transformation pipeline:

| Layer        | Purpose                             | Example                   | Typical Materialization |
| ------------ | ----------------------------------- | ------------------------- | ----------------------- |
| Staging      | Clean and standardize raw data      | `stg_customers.sql`       | View                    |
| Intermediate | Combine multiple staging models     | `int_customer_orders.sql` | View                    |
| Marts        | Business-ready tables for analytics | `dim_customers.sql`       | Table                   |

#### Domain-Based Organization

For larger projects, you might organize by business domain first, then by layer:

```
models/
├── marketing/
│   ├── staging/
│   ├── intermediate/
│   └── marts/
├── finance/
│   ├── staging/
│   ├── intermediate/
│   └── marts/
└── product/
    ├── staging/
    ├── intermediate/
    └── marts/
```

***

### Working with Tests

dbt supports two types of tests:

#### Schema Tests

These are defined in YAML files alongside your models:

```yaml
version: 2

models:
  - name: customers
    columns:
      - name: customer_id
        tests:
          - unique
          - not_null
      - name: email
        tests:
          - unique
```

#### Singular Tests

These are custom SQL queries in the `tests/` directory that should return zero rows when the test passes:

```sql
-- tests/orders_with_invalid_customer.sql
SELECT
    o.order_id
FROM {{ ref('stg_orders') }} o
LEFT JOIN {{ ref('stg_customers') }} c
    ON o.customer_id = c.customer_id
WHERE c.customer_id IS NULL
```

***

### Real-World Project Organization Example

Here's how a complete e-commerce dbt project might be organized:

```
ecommerce_dbt/
├── models/
│   ├── staging/
│   │   ├── stg_customers.sql
│   │   ├── stg_orders.sql
│   │   └── sources.yml
│   ├── intermediate/
│   │   ├── int_order_items.sql
│   │   └── int_customer_orders.sql
│   └── marts/
│       ├── core/
│       │   ├── dim_customers.sql
│       │   ├── dim_products.sql
│       │   ├── fact_orders.sql
│       │   └── schema.yml
│       └── marketing/
│           ├── customer_segmentation.sql
│           └── campaign_performance.sql
├── seeds/
│   ├── country_codes.csv
│   └── product_categories.csv
├── snapshots/
│   └── order_status_history.sql
├── macros/
│   ├── generate_schema_name.sql
│   └── cents_to_dollars.sql
├── tests/
│   └── order_price_check.sql
├── analyses/
│   └── customer_ltv.sql
├── dbt_project.yml
└── packages.yml
```

### Best Practices

* **Be consistent with naming conventions**: Use prefixes like `stg_`, `int_`, `dim_` and `fct_` to indicate model purpose
* **Document as you go**: Add descriptions in your YAML files for models and columns
* **Start simple**: Begin with a staging/marts approach and add complexity as needed
* **Group related models**: Keep related transformations close together
* **Limit cross-schema references**: Staging should only reference sources, intermediate should only reference staging, etc.
* **Use packages**: Don't reinvent common patterns when packages can help

By following a structured approach to organizing your dbt project, you'll create a more maintainable, understandable codebase that enables collaboration and scales with your team.


# Working with Sources

Learn how dbt™ sources help manage raw data tables efficiently. This guide covers defining sources in YAML, using the source() function, tracking dependencies, and ensuring data freshness.

Sources in dbt represent the raw data tables in your warehouse that serve as the foundation for your transformations. Rather than referencing raw tables directly, dbt allows you to define sources in a centralized way, improving maintainability and enabling powerful features like freshness checking.

### What Are Sources?

In dbt, **sources** represent raw data tables from external systems, such as an operational database, CRM, or third-party APIs. Instead of referencing raw tables directly in models, dbt allows you to define sources in a **centralized file** (`sources.yml`) for better organization, maintainability, and documentation.

The `sources.yml` file is a **crucial component** in dbt projects, centralizing metadata about raw data tables. This ensures consistency, maintainability, and automatic documentation.

{% hint style="info" %}
**Why Use Sources?**

* **Centralizes raw table definitions** – Avoids hardcoded table names across multiple models.
* **Improves maintainability** – If raw table locations change, you only need to update `sources.yml`.
* **Enables freshness checks** – dbt can monitor source data latency.
* **Enhances documentation** – Automatically generates lineage graphs and model dependencies.
  {% endhint %}

***

### Defining Sources in YAML

Sources are defined in `.yml` files under the `sources:` key. Here's a typical example:

```yaml
version: 2

sources:
  - name: jaffle_shop  # Logical name of the source
    database: raw      # The database where the source is stored (optional)
    schema: jaffle_shop  # Schema containing the source tables
    tables:
      - name: orders
        columns:
          - name: id
            tests:
              - unique
              - not_null
          - name: status
            tests:
              - accepted_values:
                  values: ['placed', 'shipped', 'completed', 'returned']
      - name: customers
```

In this example:

* We've defined a source named `jaffle_shop` that points to tables in the `raw.jaffle_shop` schema
* We've defined two tables: `orders` and `customers`
* We've added column-level tests to the `orders` table

***

### Using Sources in Models

Once sources are defined, you can reference them using the `source()` function in your dbt models:

```sql
-- models/staging/stg_orders.sql
SELECT
  order_id,
  customer_id,
  order_date,
  status,
  amount
FROM {{ source('jaffle_shop', 'orders') }}
```

This offers several advantages:

* **Consistency**: Source references are standardized across your project
* **Refactoring**: If a source table moves, you only need to update one place
* **Documentation**: dbt automatically builds lineage from sources to models
* **Testing**: You can apply tests to sources for early validation

***

### Best Practices for Source Organization

#### Group Related Sources

Organize sources by system or domain:

```yaml
sources:
  - name: stripe       # Payment processing
    schema: raw_stripe
    tables:
      - name: charges
      - name: customers

  - name: shopify      # E-commerce platform
    schema: raw_shopify
    tables:
      - name: orders
      - name: products
```

#### Document Your Sources

Add descriptions to help your team understand the data:

```yaml
sources:
  - name: google_analytics
    description: "Web analytics data from our marketing site"
    tables:
      - name: sessions
        description: "User sessions with UTM parameters"
        columns:
          - name: session_id
            description: "Unique identifier for the session"
```

#### Apply Tests to Sources

Find data quality issues early by testing your sources:

```yaml
sources:
  - name: crm
    tables:
      - name: customers
        columns:
          - name: customer_id
            tests:
              - unique
              - not_null
          - name: email
            tests:
              - unique
              - not_null
```

***

### Source Freshness

One of the most powerful features of sources is the ability to check data freshness - ensuring your source data is up-to-date before you build models on top of it.

#### Configuring Freshness Checks

Add a `freshness` block and specify a `loaded_at_field` in your sources definition:

```yaml
sources:
  - name: sales_data
    schema: raw_sales
    freshness:
      warn_after: {count: 12, period: hour}
      error_after: {count: 24, period: hour}
    loaded_at_field: updated_at
    tables:
      - name: transactions
```

This configuration:

* Uses the `updated_at` column to determine when data was last loaded
* Warns if data is more than 12 hours old
* Errors if data is more than 24 hours old

#### Running Freshness Checks

Check freshness with:

```bash
dbt source freshness
```

The output will show the status of each source:

```
16:35:31 | Freshness of jaffle_shop.orders: PASS (0 seconds)
16:35:32 | Freshness of jaffle_shop.customers: WARN (13 hours)
```

#### Table-Specific Freshness

You can override source-level freshness settings for specific tables:

```yaml
sources:
  - name: inventory
    freshness:
      warn_after: {count: 12, period: hour}
      error_after: {count: 24, period: hour}
    loaded_at_field: last_updated
    tables:
      - name: daily_stock
      - name: real_time_stock
        freshness:
          warn_after: {count: 15, period: minute}
          error_after: {count: 30, period: minute}
        loaded_at_field: timestamp
```

In this example, `real_time_stock` has stricter freshness requirements than other tables in the source.

***

### Advanced Source Configurations

#### Source Overrides by Environment

You can override source details for different environments by using custom schemas:

```yaml
sources:
  - name: marketing
    database: "{% if target.name == 'prod' %}analytics{% else %}raw_data{% endif %}"
    schema: "{% if target.name == 'prod' %}production{% else %}{{ target.schema }}{% endif %}"
    tables:
      - name: ad_campaigns
```

#### Filtering Source Data

For large source tables, you can define filter conditions:

```yaml
sources:
  - name: logs
    tables:
      - name: application_logs
        external:
          location: "s3://my-bucket/logs/"
          options:
            format: parquet
        freshness:
          filter: "date_column >= dateadd('day', -3, current_date)"
```

***

### Automating Source Definitions with DinoAI

With [DinoAI](/app-help/documentation/dino-ai) (Paradime's AI Agent), you can automatically generate source definitions with appropriate freshness configurations. See [step by step instructions](/app-help/documentation/dino-ai/agent-mode/use-cases/creating-sources-from-your-warehouse#step-by-step-instructions) for more details.

{% hint style="info" %}
**Benefits of automating Source definitions with DinoAI:**

✅ **Scans your data warehouse** and auto-generates the correct table definitions.\
✅ **Prevents manual errors** in source definitions.\
✅ **Keeps sources up to date** with your evolving data warehouse schema.
{% endhint %}

***

### Common Source Patterns

#### Staging Models for Sources

A common pattern is to create staging models that select from sources. These provide a clean interface between raw data and your transformations:

```sql
-- models/staging/stg_customers.sql
SELECT
  customer_id,
  first_name,
  last_name,
  email,
  created_at,
  updated_at
FROM {{ source('crm', 'customers') }}
```

#### Testing Complex Source Relationships

You can test relationships between source tables:

```yaml
sources:
  - name: application
    tables:
      - name: users
        columns:
          - name: user_id
            tests:
              - unique
              - not_null
      - name: orders
        columns:
          - name: user_id
            tests:
              - relationships:
                  to: source('application', 'users')
                  field: user_id
```

By effectively managing sources in dbt, you build a strong foundation for your analytics pipeline, making it easier to maintain, test, and document the origin of your data.


# Testing Data Quality

Discover how dbt™ testing ensures data accuracy and reliability. Learn about built-in and custom tests, schema validations, referential integrity checks, and best practices for maintaining high-qualit

Ensuring data quality is critical for any analytics pipeline. dbt provides built-in testing capabilities that help catch issues early, enforce data integrity, and maintain confidence in your transformations. This guide explains how to implement tests in your dbt project.

### Why Testing Matters in dbt

Data tests serve several essential purposes:

* **Validate assumptions** about your data
* **Catch errors** before they impact downstream consumers
* **Document expectations** about data properties
* **Ensure consistency** across transformations

Without testing, issues can creep into your data pipeline, potentially leading to incorrect business decisions or loss of trust in your analytics.

{% hint style="info" %}
**Benefits of dbt Testing**

* **Ensures Data Integrity** – Prevents duplicates, null values, and referential mismatches.
* **Validates Business Logic** – Confirms that data meets expected criteria.
* **Catches Issues Early** – Detects errors before they affect downstream analytics.
* **Automates Quality Checks** – Reduces the need for manual data validation.
* **Supports Collaboration** – Helps teams align on data expectations.
  {% endhint %}

***

### Types of dbt Tests

dbt supports two main types of tests:

#### 1. Generic Tests (Built-in)

Generic tests are reusable test definitions that can be applied to multiple models and columns. dbt includes four built-in generic tests:

| Test              | Purpose                                                  | Example Use                          |
| ----------------- | -------------------------------------------------------- | ------------------------------------ |
| `unique`          | Ensures a column has no duplicate values                 | Primary keys, email addresses        |
| `not_null`        | Ensures a column contains no NULL values                 | Required fields, join keys           |
| `accepted_values` | Validates that column values are within a specified list | Status fields, categories            |
| `relationships`   | Ensures referential integrity between tables             | Foreign keys, dimensional references |

#### 2. Singular Tests

Singular tests are custom SQL queries that define specific test logic. These are one-off tests written as SQL queries that return failing records.

***

### Adding Tests to Your Models

Tests in dbt are typically defined in YAML files alongside your models.

#### Generic Tests Example

```yaml
# models/schema.yml
version: 2

models:
  - name: customers
    columns:
      - name: customer_id
        tests:
          - unique
          - not_null
      - name: email
        tests:
          - unique
          - not_null
      - name: status
        tests:
          - accepted_values:
              values: ['active', 'inactive', 'pending']
      - name: country_id
        tests:
          - relationships:
              to: ref('countries')
              field: id
```

This YAML configuration:

* Tests that `customer_id` values are unique and not null
* Tests that `email` values are unique and not null
* Tests that `status` values are only 'active', 'inactive', or 'pending'
* Tests that each `country_id` exists in the `countries` table

#### Singular Test Example

Singular tests are SQL files in the `tests/` directory:

```sql
-- tests/assert_total_payment_amount_matches_order_amount.sql
-- This test checks that payment amounts sum to order amounts
SELECT
  orders.order_id,
  orders.amount as order_amount,
  SUM(payments.amount) as payment_amount
FROM {{ ref('orders') }}
LEFT JOIN {{ ref('payments') }} ON orders.order_id = payments.order_id
GROUP BY orders.order_id, orders.amount
HAVING ABS(orders.amount - SUM(COALESCE(payments.amount, 0))) > 0.01
```

This test identifies orders where the payment amounts don't match the order amount.

***

### Running Tests

dbt makes it easy to run tests as part of your workflow.

| Operation                    | Command                          | Description                                     |
| ---------------------------- | -------------------------------- | ----------------------------------------------- |
| Running All Tests            | `dbt test`                       | Runs all tests in your dbt project              |
| Testing Specific Models      | `dbt test --models customers`    | Runs all tests associated with a specific model |
| Running a Single Test        | `dbt test --select test_name`    | Runs a specific test by name                    |
| Testing Critical Models Only | `dbt test --select tag:critical` | Runs tests only for models tagged as 'critical' |

***

### Test Configuration Options

You can configure how tests behave using additional parameters.

#### Setting Severity Levels

Tests can be warnings instead of errors:

```yaml
models:
  - name: orders
    columns:
      - name: status
        tests:
          - accepted_values:
              values: ['completed', 'shipped', 'returned']
              severity: warn  # Won't cause pipelines to fail
```

#### Store Test Failures

Save test failures for analysis:

```yaml
models:
  - name: large_table
    columns:
      - name: id
        tests:
          - unique:
              config:
                store_failures: true  # Saves failures to a table
```

#### Limiting Failure Volume

Control how many failures are reported:

```yaml
models:
  - name: events
    columns:
      - name: event_id
        tests:
          - unique:
              config:
                limit: 100  # Only show first 100 failures
```

***

### Creating Custom Generic Tests

You can extend dbt's testing capabilities by creating custom generic tests as macros:

```sql
-- macros/test_is_valid_email.sql
{% macro test_is_valid_email(model, column_name) %}

select *
from {{ model }}
where 
    {{ column_name }} is not null 
    and not regexp_like(
        {{ column_name }}, 
        '^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\\.[a-zA-Z]{2,}$'
    )

{% endmacro %}
```

Then use it just like built-in tests:

```yaml
models:
  - name: customers
    columns:
      - name: email
        tests:
          - is_valid_email
```

***

### Test Organization Strategies

As your project grows, organizing tests becomes important:

#### Test by Business Domain

Group tests alongside the models they validate:

```
models/
├── marketing/
│   ├── schema.yml      # Contains tests for marketing models
│   ├── campaigns.sql
│   └── ad_performance.sql
└── finance/
    ├── schema.yml      # Contains tests for finance models
    ├── transactions.sql
    └── accounts.sql
```

#### Centralized Tests

Maintain all tests in a dedicated location:

```
models/
└── ...
tests/
├── generic/            # Custom generic tests
│   └── is_valid_email.sql
└── singular/           # Singular tests
    ├── marketing/
    │   └── campaign_consistency.sql
    └── finance/
        └── transaction_reconciliation.sql
```

***

### Troubleshooting Failed Tests

When tests fail, dbt provides information to help diagnose the issue:

1. **Review test SQL**: dbt outputs the SQL for the failing test
2. **Examine failing records**: Look at examples of failing records
3. **Check compiled SQL**: Review the compiled test SQL in `target/compiled/`
4. **Store failures**: Use `store_failures: true` to analyze failure patterns

***

### Test Output Examples

#### Passing Tests

```
18:42:10 | 1 of 4 START test not_null_orders_order_id ............................ [RUN]
18:42:10 | 1 of 4 PASS not_null_orders_order_id ................................. [PASS in 0.08s]
18:42:10 | 2 of 4 START test unique_orders_order_id ............................. [RUN]
18:42:10 | 2 of 4 PASS unique_orders_order_id ................................... [PASS in 0.10s]
```

#### Failing Tests

```
18:42:11 | 3 of 4 START test accepted_values_orders_status_completed__shipped__returned ... [RUN]
18:42:11 | 3 of 4 FAIL accepted_values_orders_status_completed__shipped__returned ... [FAIL in 0.15s]
18:42:11 | 4 of 4 START test relationships_orders_customer_id__customer_id__ref_customers_ ... [RUN]
18:42:11 | 4 of 4 FAIL relationships_orders_customer_id__customer_id__ref_customers_ ... [FAIL in 0.19s]
```

For failing tests, dbt shows details about the failures:

```
Failure in test relationships_orders_customer_id__customer_id__ref_customers_
Got 2 results, expected 0.

compiled SQL at target/compiled/.../relationships_orders_customer_id__customer_id__ref_customers_.sql
```

***

### Best Practices for dbt Testing

**Test Coverage Strategy**

* **Test primary keys** with `unique` and `not_null`
* **Test foreign keys** with `relationships`
* **Test categorical fields** with `accepted_values`
* **Test business logic** with singular tests
* Focus on testing critical data first

**Test Organization**

* Use a consistent naming convention for singular tests
* Group related tests together
* Document what each test validates

**Test Execution**

* Run tests before finalizing model changes
* Include tests in CI/CD pipelines
* Alert on test failures in production

**Test Maintainability**

* Prefer generic tests for common validations
* Create custom generic tests for repeated patterns
* Use macros to generate complex test logic {% endhint %}

By implementing a robust testing strategy in dbt, you can ensure your data transformations maintain high quality and reliability, building trust in your analytics data among stakeholders.


# Models and Transformations

Explore how dbt™ models streamline data transformations using SQL. Learn about model dependencies, ref(), materialization, and Jinja macros to build modular, reusable, and testable transformations.

Models are the core building blocks of your dbt project. They define the transformations that turn raw data into analytics-ready datasets using SQL. This guide covers how to create models, manage dependencies, and leverage dbt's templating capabilities.

***

### What Are Models?

In dbt, a model is a **SQL file** that defines a transformation. When you run dbt, it compiles these SQL files into executable queries and runs them against your data warehouse, creating views or tables.

Models serve three key purposes:

1. **Transform data** into useful analytics structures
2. **Document transformations** with clear SQL
3. **Create dependencies** between different data assets

Each model typically results in a single table or view in your data warehouse.

***

### Creating Your First Model

A model is simply a `.sql` file in your `models/` directory. Let's start with a basic example:

```sql
-- models/staging/stg_customers.sql
SELECT
  id as customer_id,
  first_name,
  last_name,
  email,
  date_joined
FROM {{ source('jaffle_shop', 'customers') }}
```

When you run `dbt run`, this SQL gets compiled and executed in your data warehouse, creating a view called `stg_customers` with transformed customer data.

***

### Using Common Table Expressions (CTEs)

CTEs make your models more readable and maintainable by breaking complex queries into logical building blocks. They're a powerful way to structure your transformations:

```sql
-- models/marts/customer_orders.sql
WITH customers AS (
    SELECT * FROM {{ ref('stg_customers') }}
),

orders AS (
    SELECT * FROM {{ ref('stg_orders') }}
),

customer_orders AS (
    SELECT
        customer_id,
        COUNT(order_id) as order_count,
        SUM(amount) as total_spent
    FROM orders
    GROUP BY customer_id
)

SELECT
    customers.customer_id,
    customers.first_name,
    customers.last_name,
    customers.email,
    COALESCE(customer_orders.order_count, 0) as order_count,
    COALESCE(customer_orders.total_spent, 0) as total_spent
FROM customers
LEFT JOIN customer_orders USING (customer_id)
```

CTEs offer several benefits:

* Improve readability by breaking complex logic into named sections
* Allow you to reuse intermediate calculations
* Make troubleshooting easier by separating transformation steps

***

### Model Dependencies with ref()

The `ref()` function is one of dbt's most powerful features. It allows you to reference other models, automatically creating dependencies:

```sql
-- models/marts/customer_lifetime_value.sql
WITH customer_orders AS (
    SELECT * FROM {{ ref('customer_orders') }}
)

SELECT
    customer_id,
    total_spent,
    total_spent * 0.15 as estimated_future_value,
    total_spent * 1.15 as lifetime_value
FROM customer_orders
```

When you use `ref()`:

1. dbt automatically determines the correct schema and table name
2. dbt builds a dependency graph, ensuring models run in the correct order
3. dbt creates lineage documentation for your project

A key benefit is that if you rename models or change schemas, dbt handles all the references for you.

***

### Model Configuration with config()

The `config()` function lets you control how a model is materialized and other settings:

```sql
-- models/marts/large_summary_table.sql
{{ 
  config(
    materialized='table',
    sort='date_day',
    dist='customer_id'
  ) 
}}

SELECT
  date_trunc('day', created_at) as date_day,
  customer_id,
  sum(amount) as total_amount
FROM {{ ref('stg_orders') }}
GROUP BY 1, 2
```

Common configuration options include:

* `materialized`: How the model should be created ('view', 'table', 'incremental', 'ephemeral')
* `schema`: Which schema the model should be created in
* `tags`: Labels to organize and select models
* Database-specific options (like `sort`, `dist`, `cluster_by`, etc.)

***

### Using Jinja for Dynamic SQL

Jinja is a templating language that allows you to generate dynamic SQL. dbt uses Jinja to make your transformations more flexible and reusable.

#### Conditional Logic

Use if/else statements to adapt your SQL based on conditions:

```sql
SELECT
  order_id,
  order_date,
  {% if target.name == 'prod' %}
    amount
  {% else %}
    amount * 100 as amount_in_cents
  {% endif %}
FROM {{ ref('stg_orders') }}
```

#### Looping

Generate repetitive SQL using for loops:

```sql
SELECT
  order_id,
  {% for i in range(1, 5) %}
    item_{{ i }}_id,
    item_{{ i }}_quantity,
    {% if not loop.last %},{% endif %}
  {% endfor %}
FROM {{ ref('stg_order_items') }}
```

#### Variables

Use variables to make your models configurable:

```sql
-- Using a variable defined in dbt_project.yml or passed via --vars
SELECT *
FROM {{ ref('stg_orders') }}
WHERE order_date >= '{{ var("start_date", "2020-01-01") }}'
```

***

### Model Organization Best Practices

Organize your models to reflect their purpose in your analytics pipeline:

#### Staging Models

Staging models clean and standardize source data:

* Naming and datatype standardization
* Simple filtering
* One-to-one relationship with source tables
* Typically materialized as views

```sql
-- models/staging/stg_customers.sql
SELECT
  id as customer_id,
  first_name,
  last_name,
  email,
  -- Convert to ISO date format
  PARSE_DATE('%Y-%m-%d', date_joined) as date_joined
FROM {{ source('jaffle_shop', 'customers') }}
WHERE id IS NOT NULL
```

#### Intermediate Models

Intermediate models combine and transform staging models:

* Join related data sources
* Apply business logic
* Create reusable building blocks
* Typically materialized as views

```sql
-- models/intermediate/int_customer_orders.sql
SELECT
  o.order_id,
  o.customer_id,
  c.email,
  o.order_date,
  o.status,
  o.amount
FROM {{ ref('stg_orders') }} o
JOIN {{ ref('stg_customers') }} c ON o.customer_id = c.customer_id
```

#### Mart Models

Mart models prepare data for business consumption:

* Oriented around business entities (customers, products, etc.)
* Optimized for specific use cases
* Include calculated metrics
* Often materialized as tables for performance

```sql
-- models/marts/finance/order_payment_summary.sql
{{
  config(
    materialized='table'
  )
}}

SELECT
  date_trunc('month', o.order_date) as order_month,
  p.payment_method,
  count(distinct o.order_id) as order_count,
  sum(o.amount) as total_amount
FROM {{ ref('int_customer_orders') }} o
JOIN {{ ref('stg_payments') }} p ON o.order_id = p.order_id
GROUP BY 1, 2
```

***

### Troubleshooting Models

When your models have issues, use these strategies to troubleshoot:

#### Compiling Without Running

Use `dbt compile` to see the generated SQL without running it:

```bash
dbt compile --models customer_lifetime_value
```

Then check the compiled SQL in `target/compiled/[project_name]/models/...`

#### Execute Specific Models

Run only the model you're working on:

```bash
dbt run --models staging.stg_customers
```

Or run a model and everything that depends on it:

```bash
dbt run --models +customer_lifetime_value
```

#### Check Logs

Detailed logs are available in the `logs/` directory and often contain helpful error information.

***

### Advanced Model Techniques

#### Custom Schemas

Generate custom schemas to separate models by team or environment:

```sql
{{ 
  config(
    schema='marketing_' ~ target.name
  ) 
}}

SELECT * FROM {{ ref('stg_marketing_campaigns') }}
```

#### Post-Hooks

Execute SQL after a model is created, such as granting permissions:

```sql
{{ 
  config(
    post_hook='GRANT SELECT ON {{ this }} TO ROLE analytics_readers'
  ) 
}}

SELECT * FROM {{ ref('stg_customers') }}
```

#### Documentation in Models

Add descriptions to your models using YAML files:

```yaml
version: 2

models:
  - name: customer_orders
    description: "One record per customer with order summary data"
    columns:
      - name: customer_id
        description: "The primary key of the customers table"
      - name: order_count
        description: "Count of orders placed by this customer"
        tests:
          - not_null
      - name: total_spent
        description: "Total amount spent on all orders"
```

These descriptions will appear in your automatically generated documentation.

By mastering models and transformations in dbt, you can build a reliable, maintainable analytics pipeline that transforms raw data into valuable business insights.


# Configuring your dbt™ Project

Setting up a well-structured **dbt™ project** is essential for scalability, maintainability, and efficiency. This section covers the core configurations needed to define sources, manage transformations, and ensure data quality in your dbt™ workflows.

Whether you're starting from scratch or optimizing an existing setup, these guides will help you configure your dbt™ project effectively.

### Key Areas of Configuration

<table data-view="cards"><thead><tr><th></th><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th><th data-hidden data-card-cover data-type="files"></th></tr></thead><tbody><tr><td>📄 <strong>Setting Up Your dbt_project.yml</strong></td><td>Define project-wide configurations, including materializations, model directories, and environment settings.</td><td></td><td><a href="/pages/szCNwMvK0EcMzWcFj8Wi">/pages/szCNwMvK0EcMzWcFj8Wi</a></td><td></td></tr><tr><td>📊 <strong>Defining Your Sources</strong></td><td>Use <code>sources.yml</code> to document and reference external data sources in your transformations.</td><td></td><td><a href="/pages/kqwNSf5WeloUKjfDUU2r">/pages/kqwNSf5WeloUKjfDUU2r</a></td><td></td></tr><tr><td>🔄 <strong>Testing Source Freshness</strong></td><td>Ensure your raw data is up to date with automatic freshness checks in dbt™.</td><td></td><td><a href="/pages/f7WZmJpI4tZIE3vSU2Uk">/pages/f7WZmJpI4tZIE3vSU2Uk</a></td><td></td></tr><tr><td>🏷️ <strong>Working with Tags in Your dbt™ Project</strong></td><td>Organize and selectively run models using tags for better project structure and workflow control.</td><td></td><td><a href="/pages/634S2oxI9LrvAiufTxR4">/pages/634S2oxI9LrvAiufTxR4</a></td><td></td></tr><tr><td><strong>🧪 Unit Testing</strong></td><td>Validate your SQL transformation logic with controlled input data before deploying models.</td><td></td><td><a href="/pages/sJ9GhmaddDQ1J3M5woD4">/pages/sJ9GhmaddDQ1J3M5woD4</a></td><td></td></tr></tbody></table>

{% hint style="info" %}
Prefer hands-on learning? Check out our [**Paradime 101 Guide**](/app-help/guides/paradime-101) for a step-by-step, interactive way to learn dbt™ and analytics engineering best practices—all for free.
{% endhint %}


# Setting up your dbt\_project.yml

Explains the crucial dbt\_project.yml configuration file, including its purpose, key components, and best practices for maintaining an organized and evolving dbt project.

The `dbt_project.yml` file is the **core configuration file** for any dbt project. It defines essential settings such as the project name, version, and model configurations, ensuring your project runs correctly and consistently.

***

### Why dbt\_project.yml Matters

The `dbt_project.yml` file serves several important functions:

* Identifies the root of your dbt project
* Configures project-wide settings
* Sets default materializations for your models
* Defines model-specific configurations
* Controls directory paths and behaviors

A well-configured project file ensures consistent behavior across environments and team members.

***

### Core Components of dbt\_project.yml

Here's a breakdown of the key sections and their purposes:

#### Project Metadata

```yaml
name: 'my_dbt_project'  # The unique name of your dbt project
version: '1.0.0'        # Optional versioning for project tracking
config-version: 2       # The version of dbt's config schema
```

This section defines:

* `name`: A unique identifier for your project (used in compiled SQL)
* `version`: Optional versioning for tracking project changes
* `config-version`: The version of dbt's configuration schema (should be 2 for current projects)

#### Profile Configuration

```yaml
profile: 'my_profile'  # Specifies the profile to use from profiles.yml
```

This tells dbt which profile to use from your `profiles.yml` file. Profiles define database connections and credentials.

| Setting   | Purpose                                   | Example                          |
| --------- | ----------------------------------------- | -------------------------------- |
| `profile` | Specifies which connection profile to use | `profile: 'snowflake_analytics'` |

#### Directory Paths

```yaml
# Paths for different dbt components
model-paths: ["models"]
analysis-paths: ["analyses"]
test-paths: ["tests"]
seed-paths: ["seeds"]
macro-paths: ["macros"]
snapshot-paths: ["snapshots"]
```

These settings define where dbt should look for different types of files:

| Path Setting     | Default         | Purpose                               |
| ---------------- | --------------- | ------------------------------------- |
| `model-paths`    | `["models"]`    | Where your SQL models are stored      |
| `seed-paths`     | `["seeds"]`     | Where your CSV files are stored       |
| `test-paths`     | `["tests"]`     | Where singular tests are stored       |
| `analysis-paths` | `["analyses"]`  | Where analytical queries are stored   |
| `macro-paths`    | `["macros"]`    | Where macros are stored               |
| `snapshot-paths` | `["snapshots"]` | Where snapshot definitions are stored |

#### Model Configuration

This section defines how your models should be materialized and configured:

```yaml
models:
  my_dbt_project:  # Must match your project name
    +materialized: view   # Default materialization for all models
    
    # Configure specific directories
    staging:
      +materialized: view  # Staging models as views
    
    marts:
      +materialized: table # Mart models as tables
      
      # Configure specific subdirectories
      marketing:
        +schema: marketing_schema  # Custom schema
```

Key points about model configuration:

* Configuration is hierarchical - lower levels inherit from higher levels
* The top-level project name must match your `name` value
* The `+` prefix indicates a dbt configuration property
* You can override configurations at any level

#### Seed Configuration

For controlling how CSV files are loaded into your database:

```yaml
seeds:
  my_dbt_project:
    +schema: raw_data     # Default schema for seed files
    +quote_columns: false # Whether to quote column names
    
    # Configuration for specific seeds
    country_codes:
      +column_types:
        country_code: varchar(2)
```

#### Variables

Define project-wide variables that can be used in models:

```yaml
vars:
  start_date: '2020-01-01'  # Available as {{ var('start_date') }}
  countries: ['US', 'CA', 'UK']
  # Environment-specific variables
  dev:
    debug_mode: true
  prod:
    debug_mode: false
```

#### On-Run Hooks

Execute SQL before or after your dbt runs:

```yaml
on-run-start:
  - "create schema if not exists {{ target.schema }}_staging"
  
on-run-end:
  - "grant usage on schema {{ target.schema }} to role reporter"
  - "grant select on all tables in schema {{ target.schema }} to role reporter"
```

#### Cleaning Up Artifacts

Define which directories should be cleaned by `dbt clean`:

```yaml
clean-targets:
  - "target"
  - "dbt_packages"
  - "logs"
```

***

### Complete Example

Here's a complete example of a `dbt_project.yml` file:

```yaml
name: 'ecommerce'
version: '1.0.0'
config-version: 2

profile: 'snowflake_analytics'

model-paths: ["models"]
seed-paths: ["seeds"]
test-paths: ["tests"]
analysis-paths: ["analyses"]
macro-paths: ["macros"]
snapshot-paths: ["snapshots"]
docs-paths: ["docs"]

target-path: "target"
clean-targets:
  - "target"
  - "dbt_packages"
  - "logs"

vars:
  start_date: '2020-01-01'
  include_test_accounts: false

models:
  ecommerce:
    +materialized: view
    
    staging:
      +materialized: view
      +schema: staging
      
    intermediate:
      +materialized: view
      +schema: intermediate
    
    marts:
      +materialized: table
      +schema: analytics
      
      finance:
        +schema: analytics_finance
        +tags: ["finance", "daily"]
      
      marketing:
        +schema: analytics_marketing
        +tags: ["marketing"]

seeds:
  ecommerce:
    +schema: reference_data

snapshots:
  ecommerce:
    +target_schema: snapshots

on-run-end:
  - "grant select on all tables in schema {{ target.schema }} to role analyst"
```

***

### Best Practices for dbt\_project.yml

| Category                           | Best Practices                                                                                                                                                                                                 |
| ---------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Use Meaningful Names and Structure | <p>✅ Group models logically by function or business domain<br>✅ Use consistent naming patterns for schemas<br>✅ Document non-obvious configurations with comments</p>                                          |
| Set Sensible Defaults              | <p>✅ Define default materializations for different model types<br>✅ Use views for staging/intermediate models and tables for final models<br>✅ Configure schemas to match your data warehouse organization</p> |
| Optimize for Team Collaboration    | <p>✅ Use environment-specific variables where needed<br>✅ Set appropriate permissions with on-run hooks<br>✅ Document variables and their purposes</p>                                                         |
| Maintain and Evolve                | <p>✅ Review your project configuration regularly<br>✅ Update as your project grows and changes<br>✅ Document changes to configuration in version control</p>                                                   |

***

### Common Issues and Solutions

| Issue                           | Solution                                              |
| ------------------------------- | ----------------------------------------------------- |
| Models building in wrong schema | Check schema configuration and target profile         |
| Incorrect materialization       | Verify hierarchy of materialization settings          |
| Variable not available          | Ensure variable is defined at the correct level       |
| Path not found                  | Verify directory paths match actual project structure |

Your `dbt_project.yml` file is a living document that will evolve with your project. Taking the time to configure it correctly will lead to a more maintainable and consistent dbt implementation.


# Defining Your Sources in sources.yml

Covers how to properly define and use sources in dbt™ by configuring the    \`sources.yml\` file, referencing sources in models, and ensuring data quality.

Sources in dbt represent the raw data tables in your data warehouse. Defining sources in a `sources.yml` file allows you to centralize table references, enabling cleaner code, testing, and documentation.

***

### What is sources.yml?

The `sources.yml` file is where you define the external data tables that your dbt models will transform. This file centralizes information about your raw data sources, making it easier to:

* Reference raw tables consistently throughout your project
* Test source data for quality and freshness
* Document your data pipeline from beginning to end
* Create clear lineage visualization in dbt docs

***

### Basic Structure

A basic `sources.yml` file follows this structure:

```yaml
version: 2

sources:
  - name: jaffle_shop  # Source system name
    database: raw      # Optional: database where source is stored
    schema: jaffle_shop  # Schema containing the source tables
    tables:
      - name: orders   # Table name as it exists in the database
      - name: customers
```

This defines a source called `jaffle_shop` with two tables: `orders` and `customers`.

***

### Required and Optional Fields

| Field         | Required? | Description                                                           |
| ------------- | --------- | --------------------------------------------------------------------- |
| `name`        | Required  | Logical name for the source (used in the `source()` function)         |
| `schema`      | Required  | Database schema where the tables exist                                |
| `database`    | Optional  | Database where the schema exists (if different from target database)  |
| `tables`      | Required  | List of tables in this source                                         |
| `description` | Optional  | Description of the source for documentation                           |
| `loader`      | Optional  | Information about the tool loading this data (e.g., Fivetran, Stitch) |
| `freshness`   | Optional  | Configuration for freshness checks                                    |
| `quoting`     | Optional  | Settings for quoting identifiers                                      |

***

### Detailed Source Configuration

Here's a more complete example with additional configurations:

```yaml
version: 2

sources:
  - name: stripe
    description: "Payment data from Stripe, loaded by Fivetran connector"
    database: raw_data
    schema: stripe
    loader: fivetran
    loaded_at_field: _fivetran_synced
    
    # Source-level freshness (can be overridden at table level)
    freshness:
      warn_after: {count: 12, period: hour}
      error_after: {count: 24, period: hour}
    
    tables:
      - name: charges
        description: "Credit card charges, including status and amount"
        columns:
          - name: id
            description: "Primary key - the Stripe charge ID"
            tests:
              - unique
              - not_null
          - name: amount
            description: "Amount in cents"
            tests:
              - not_null
      
      - name: customers
        description: "Customer information from Stripe"
        # Override source-level freshness for this table
        freshness:
          warn_after: {count: 24, period: hour}
          error_after: {count: 48, period: hour}
```

***

### Referencing Sources in Models

Once defined, you can reference sources in your models using the `source()` function:

```sql
-- models/staging/stg_stripe_charges.sql
SELECT
  id as charge_id,
  customer_id,
  amount / 100.0 as amount_usd,  -- Convert cents to dollars
  status,
  created as created_at
FROM {{ source('stripe', 'charges') }}
```

The `source()` function takes two arguments:

1. The source name (defined as `name` in `sources.yml`)
2. The table name (defined under `tables` in `sources.yml`)

***

### Testing Sources

You can apply tests to your source tables just like you would with models:

```yaml
sources:
  - name: crm
    tables:
      - name: customers
        columns:
          - name: customer_id
            tests:
              - unique
              - not_null
          - name: email
            tests:
              - unique
              - not_null
```

This allows you to verify data quality at the entry point to your dbt pipeline.

***

### Source Freshness

One of the most powerful features of sources is the ability to check data freshness - ensuring your source data is up-to-date.

#### Configuring Freshness Checks

To set up freshness checks, you need:

1. A column in your source table that indicates when a record was last updated (`loaded_at_field`)
2. Freshness threshold configurations

```yaml
sources:
  - name: stripe
    loaded_at_field: _etl_loaded_at
    freshness:
      warn_after: {count: 12, period: hour}
      error_after: {count: 24, period: hour}
    tables:
      - name: charges
      - name: customers
        # Override for specific table
        freshness:
          warn_after: {count: 36, period: hour}
```

{% hint style="info" %}
**Valid Time Periods for Freshness Checks**

When configuring source freshness thresholds, you must specify both a count and a time period. The time period can be:

* `minute` - For data that updates very frequently (e.g., real-time systems)
* `hour` - For data that updates throughout the day (e.g., transaction systems)
* `day` - For data that updates daily or less frequently (e.g., batch uploads)

For example: `warn_after: {count: 6, period: hour}` would warn if data is more than 6 hours old.
{% endhint %}

#### Running Freshness Checks

Execute freshness checks with:

```bash
dbt source freshness
```

The output indicates which sources are fresh, stale (warning), or too old (error).

***

### Advanced Source Configurations

#### Dynamic Schema Resolution

You can use Jinja in your source definitions to handle environments:

```yaml
sources:
  - name: marketing
    schema: "{% if target.name == 'prod' %}marketing_prod{% else %}marketing_{{ target.name }}{% endif %}"
    tables:
      - name: campaigns
```

This creates different schemas depending on your target environment.

#### Quoting Configuration

Control identifier quoting:

```yaml
sources:
  - name: salesforce
    quoting:
      database: true
      schema: true
      identifier: false
    tables:
      - name: Contacts
        quoting:
          identifier: true  # Override source-level setting
```

#### External Tables

For data sources like S3 or GCS:

```yaml
sources:
  - name: clickstream
    tables:
      - name: raw_events
        external:
          location: "s3://my-bucket/clickstream/events/"
          file_format: "parquet"
```

***

### Best Practices for Source Configuration

| Practice                             | Description                                                      |
| ------------------------------------ | ---------------------------------------------------------------- |
| Group logically related sources      | Organize sources by system or domain (e.g., CRM, ERP, Marketing) |
| Document sources thoroughly          | Add descriptions for sources, tables, and columns                |
| Test source data                     | Apply tests to key columns to catch issues early                 |
| Set appropriate freshness thresholds | Define freshness based on business needs and data load frequency |
| Use consistent naming patterns       | Establish a convention for source names and stick to it          |
| Consider file organization           | Place source definitions close to the models that use them       |

***

### Organizing source.yml Files

There are several approaches to organizing your source definitions:

#### Single File Approach

Good for smaller projects:

```
models/
  └── sources.yml    # All sources defined in one file
```

#### Source-by-Source Approach

Better for larger projects:

```
models/
  └── sources/
      ├── salesforce_sources.yml
      ├── stripe_sources.yml
      └── google_analytics_sources.yml
```

#### Alongside Related Models

Place source definitions with the models that use them:

```
models/
  ├── staging/
  │   ├── stripe/
  │   │   ├── stripe_sources.yml
  │   │   ├── stg_stripe_charges.sql
  │   │   └── stg_stripe_customers.sql
  │   └── salesforce/
  │       ├── salesforce_sources.yml
  │       └── stg_salesforce_contacts.sql
```

***

{% hint style="info" %}

#### Automating Source Generation with DinoAI

With [DinoAI](/app-help/documentation/dino-ai) (Paradime's AI Agent), you can automatically generate source definitions with appropriate freshness configurations. See [step by step instructions](/app-help/documentation/dino-ai/agent-mode/use-cases/creating-sources-from-your-warehouse#step-by-step-instructions) for more details.
{% endhint %}


# Testing Source Freshness

Learn how to configure and run source freshness checks in dbt™, ensuring your data is up-to-date and meets SLAs. Covers setup, optimization, and best practices for maintaining reliable data pipelines.

Ensuring your source data is up-to-date is critical for reliable analytics. dbt's source freshness checks allow you to monitor and verify when data was last loaded, helping your team maintain data quality and meet service level agreements (SLAs).

### Why Test Source Freshness?

Source freshness checks serve several important purposes:

| Purpose                     | Description                                                              |
| --------------------------- | ------------------------------------------------------------------------ |
| Detect stale data           | Identify when source data hasn't been updated within expected timeframes |
| Enforce SLAs                | Ensure that data meets freshness requirements for business operations    |
| Prevent incorrect analytics | Avoid building reports on outdated information                           |
| Monitor pipeline health     | Get early warnings when data ingestion processes fail                    |

### How Source Freshness Works

dbt's source freshness checking works by:

1. Examining a timestamp column in your source tables
2. Comparing that timestamp to the current time
3. Evaluating the difference against your defined thresholds

{% hint style="info" %}
**Example Freshness Check Failures**

When data exceeds your freshness thresholds, dbt will report failures like these:

***

```bash
03:15:22 | 1 of 3 WARN freshness of jaffle_shop.orders ........................ [WARN in 1.42s]
03:15:22 | WARN: Source jaffle_shop.orders (model: sources.yml) - Missing 15.3 hours of data. Latest record loaded at 2023-09-15 12:01:32, expected within 12 hours

03:15:24 | 2 of 3 ERROR freshness of stripe.transactions ...................... [ERROR in 2.01s]
03:15:24 | ERROR: Source stripe.transactions (model: stripe_sources.yml) - Missing 28.5 hours of data. Latest record loaded at 2023-09-14 22:45:11, expected within 24 hours
```

***

When data exceeds your freshness thresholds, dbt will report failures like these:

The warnings and errors include the source name, how much data is missing, the timestamp of the most recent record, and your configured threshold.
{% endhint %}

***

### Configuring Freshness Checks

To enable freshness checks, you need to add two key elements to your `sources.yml` file:

1. A `loaded_at_field` that identifies the timestamp column
2. A `freshness` configuration that defines your thresholds

Here's a basic example:

```yaml
version: 2

sources:
  - name: jaffle_shop
    database: raw  
    freshness:
      warn_after: {count: 12, period: hour}  # Warn if data is over 12 hours old
      error_after: {count: 24, period: hour}  # Error if data is over 24 hours old
    loaded_at_field: _etl_loaded_at  # Column storing the last update timestamp
    tables:
      - name: raw_orders
      - name: raw_customers  # Inherits default freshness settings
```

#### Understanding the Configuration

* `loaded_at_field`: The column containing the timestamp when data was loaded
* `warn_after`: When to issue a warning about stale data
* `error_after`: When to report an error about stale data (typically more lenient than warnings)

{% hint style="info" %}
**Valid Time Periods for Freshness Checks**

When configuring source freshness thresholds, you must specify both a count and a time period. The time period can be:

* `minute` - For data that updates very frequently (e.g., real-time systems)
* `hour` - For data that updates throughout the day (e.g., transaction systems)
* `day` - For data that updates daily or less frequently (e.g., batch uploads)

For example: `warn_after: {count: 6, period: hour}` would warn if data is more than 6 hours old.
{% endhint %}

***

### Table-Specific Freshness Settings

You can override source-level freshness settings for specific tables:

```yaml
sources:
  - name: jaffle_shop
    freshness:
      warn_after: {count: 12, period: hour}
      error_after: {count: 24, period: hour}
    loaded_at_field: _etl_loaded_at
    tables:
      - name: raw_orders
        freshness:  
          warn_after: {count: 6, period: hour}  # More strict for orders table
          error_after: {count: 12, period: hour}

      - name: raw_customers  # Inherits source's default freshness

      - name: raw_product_skus
        freshness: null  # Disable freshness checks for this table
```

***

### Running Freshness Checks

To check the freshness of your source data, run:

```bash
dbt source freshness
```

This command:

1. Queries each source table for the most recent `loaded_at_field` value
2. Compares that timestamp against the current time
3. Evaluates it against your thresholds
4. Reports success, warning, or error for each source

{% hint style="info" %}
**Where Freshness Check Results Appear**

When you run `dbt source freshness`, the results appear in multiple places:

1. **Terminal Output**: Warnings and errors are displayed in your command line interface immediately when running the command
2. **Log Files**: Results are written to the log files in your project's `logs/` directory
3. **Artifacts**: Detailed results are stored in the `target/sources.json` file
4. **dbt Cloud** (if using): Results appear in the run history and can trigger notifications
5. **Paradime Interface** (if using): Results are displayed in the Paradime UI with history tracking

These results can be consumed by monitoring tools, notification systems, or dashboards to provide visibility into your data pipeline health.
{% endhint %}

#### Example Output

```
03:33:31 | Concurrency: 1 threads (target='dev')
03:33:31 | 
03:33:31 | 1 of 2 START freshness of jaffle_shop.raw_orders ................... [RUN]
03:33:32 | 1 of 2 WARN freshness of jaffle_shop.raw_orders .................... [WARN in 0.98s]
03:33:32 | 2 of 2 START freshness of jaffle_shop.raw_customers ................ [RUN]
03:33:33 | 2 of 2 PASS freshness of jaffle_shop.raw_customers ................. [PASS in 0.82s]
03:33:33 | 
03:33:33 | Finished running 2 source freshness checks in 1.99s.
```

***

### Integrating Freshness Checks into Workflows

Freshness checks can be integrated into your data pipeline in several ways:

#### Scheduled Checks

Run freshness checks on a regular schedule to proactively monitor your data:

```bash
# Example cron job to check freshness every hour
0 * * * * cd /path/to/project && dbt source freshness
```

#### Pre-run Validation

Execute freshness checks before running models to prevent building on stale data:

```bash
#!/bin/bash
# Exit with error if any source freshness check fails
dbt source freshness || exit 1
# Only run models if freshness checks pass
dbt run
```

#### CI/CD Pipeline Integration

Include freshness checks in your CI/CD workflows:

```yaml
# Example GitHub Actions workflow step
- name: Check source freshness
  run: dbt source freshness --target prod
```

***

### Optimizing Freshness Checks for Large Tables

For very large datasets, running freshness checks can be resource-intensive. You can optimize them using filters:

```yaml
sources:
  - name: analytics
    tables:
      - name: events
        loaded_at_field: created_at
        freshness:
          warn_after: {count: 1, period: hour}
          error_after: {count: 6, period: hour}
          filter: "date_trunc('day', created_at) >= dateadd('day', -3, current_date)"
```

The `filter` clause limits the rows that dbt examines when checking freshness, improving performance for large tables.

***

### Handling Different Data Loading Patterns

Different sources may have various data loading patterns that affect how you configure freshness:

| Loading Pattern      | Freshness Strategy      | Example                                   |
| -------------------- | ----------------------- | ----------------------------------------- |
| Real-time streaming  | Short windows (minutes) | `warn_after: {count: 15, period: minute}` |
| Hourly batch updates | Medium windows (hours)  | `warn_after: {count: 2, period: hour}`    |
| Daily ETL jobs       | Longer windows (days)   | `warn_after: {count: 1, period: day}`     |
| Weekly data delivery | Extended windows        | `warn_after: {count: 8, period: day}`     |

### Troubleshooting Freshness Issues

If your freshness checks are failing, consider these common issues:

1. **Incorrect timestamp column**: Verify that `loaded_at_field` is the right column
2. **Timezone differences**: Check if there are timezone discrepancies between source timestamps and dbt
3. **Data loading failures**: Investigate upstream ETL/ELT processes
4. **Unrealistic expectations**: Adjust thresholds to match actual data loading patterns

***

### Best Practices

| Practice                 | Description                                                                                                       |
| ------------------------ | ----------------------------------------------------------------------------------------------------------------- |
| Set realistic thresholds | Align freshness requirements with business needs and actual data load frequencies                                 |
| Use appropriate column   | Choose a column that truly represents when data was last updated (ETL timestamp preferred over source timestamps) |
| Monitor trends           | Track freshness over time to identify deteriorating pipeline performance                                          |
| Disable when appropriate | Use `freshness: null` for static reference tables that rarely change                                              |
| Document expectations    | Include freshness SLAs in your data documentation                                                                 |

***

{% hint style="info" %}

#### Automating Source Generation with DinoAI

With [DinoAI](/app-help/documentation/dino-ai) (Paradime's AI Agent), you can automatically generate source definitions with appropriate freshness configurations. See [step by step instructions](/app-help/documentation/dino-ai/agent-mode/use-cases/creating-sources-from-your-warehouse#step-by-step-instructions) for more details.
{% endhint %}


# Unit Testing

Learn how to use unit tests in dbt™ to validate transformation logic, prevent errors,    and improve code reliability before deploying to production.

Unit testing in dbt allows you to validate your SQL transformation logic using controlled input data. Unlike traditional data tests that verify the quality of existing data, unit tests help you catch logical errors during development, bringing the test-driven development approach to data transformations.

***

### Understanding dbt Unit Tests

Unit tests in dbt help you verify that your transformations produce expected outputs given specific input data. This ensures your business logic is correct before you deploy it to production.

{% hint style="info" %}
Unit testing is available in dbt Core v1.8+ and above
{% endhint %}

***

### Why Unit Testing Matters

Traditional data tests (like `not_null`, `unique`) validate the quality of data that already exists. Unit tests serve a different and complementary purpose:

| Unit Tests                          | Data Tests                       |
| ----------------------------------- | -------------------------------- |
| Validate transformation logic       | Validate data quality            |
| Run before building models          | Run after models are built       |
| Use controlled test data            | Use actual production data       |
| Focus on business logic correctness | Focus on data integrity          |
| Help with test-driven development   | Help with data quality assurance |

Unit tests provide several benefits:

* **Test before building**: Validate logic without materializing models
* **Verify transformations**: Ensure your SQL logic handles edge cases correctly
* **Support test-driven development**: Write tests first, then implement the model
* **Catch errors early**: Find bugs before they reach production
* **Improve code reliability**: Maintain confidence during refactoring

***

### When to Use Unit Tests

Unit tests are particularly valuable when your models include:

* **Complex transformations**: Regular expressions, date calculations, window functions
* **Critical business logic**: Calculations that impact important metrics
* **Known edge cases**: Scenarios that have caused issues in the past
* **Models undergoing refactoring**: Changes to existing transformation logic
* **Frequently used models**: Core models that many others depend on

***

### Unit Test Structure

A dbt unit test consists of three essential parts:

1. **Mock inputs**: Sample data for source tables and referenced models
2. **Model under test**: The model whose logic you want to validate
3. **Expected outputs**: The exact results you expect after transformation

#### Basic Example

Here's a simple unit test defined in YAML:

```yaml
unit_tests:
  - name: test_orders_status_counts
    model: order_status_summary
    given:
      # Define input data for the model's dependencies
      - input: ref('stg_orders')
        rows:
          - {order_id: 1, status: 'completed'}
          - {order_id: 2, status: 'pending'}
          - {order_id: 3, status: 'completed'}
          - {order_id: 4, status: 'shipped'}
    expect:
      # Define expected output data
      rows:
        - {status: 'completed', count: 2}
        - {status: 'pending', count: 1}
        - {status: 'shipped', count: 1}
```

***

### Creating Unit Tests

Unit tests are defined in YAML files within your models directory, typically alongside the model they're testing.

#### Defining Input Data

Mock input data can be provided in several formats:

**Dictionary Format (default)**

```yaml
given:
  - input: ref('stg_customers')
    rows:
      - {customer_id: 1, name: 'John Doe', status: 'active'}
      - {customer_id: 2, name: 'Jane Smith', status: 'inactive'}
```

**CSV Format**

```yaml
given:
  - input: ref('stg_customers')
    format: csv
    rows: |
      customer_id,name,status
      1,"John Doe",active
      2,"Jane Smith",inactive
```

**SQL Format**

```yaml
given:
  - input: ref('stg_customers')
    format: sql
    rows: |
      select 1 as customer_id, 'John Doe' as name, 'active' as status
      union all
      select 2 as customer_id, 'Jane Smith' as name, 'inactive' as status
```

#### Defining Expected Output

You can specify expected outputs in different ways:

**Rows (most common)**

```yaml
expect:
  rows:
    - {customer_id: 1, status: 'active'}
    - {customer_id: 2, status: 'inactive'}
```

**Column Values**

```yaml
expect:
  columns:
    - name: status
      values: ['active', 'inactive']
```

**Row Count**

```yaml
expect:
  row_count: 2
```

### Example: Testing a Customer Classification Model

Let's see a complete example for a model that classifies customers based on spending:

**The model being tested:**

```sql
-- models/customer_segments.sql
SELECT
  customer_id,
  name,
  total_spend,
  CASE
    WHEN total_spend >= 1000 THEN 'high'
    WHEN total_spend >= 500 THEN 'medium'
    ELSE 'low'
  END as spending_segment
FROM {{ ref('stg_customers') }}
```

**The unit test:**

```yaml
# models/customer_segments_tests.yml
unit_tests:
  - name: test_customer_segments_classification
    model: customer_segments
    given:
      - input: ref('stg_customers')
        rows:
          - {customer_id: 1, name: 'Customer A', total_spend: 1200}
          - {customer_id: 2, name: 'Customer B', total_spend: 750} 
          - {customer_id: 3, name: 'Customer C', total_spend: 300}
    expect:
      rows:
        - {customer_id: 1, name: 'Customer A', total_spend: 1200, spending_segment: 'high'}
        - {customer_id: 2, name: 'Customer B', total_spend: 750, spending_segment: 'medium'}
        - {customer_id: 3, name: 'Customer C', total_spend: 300, spending_segment: 'low'}
```

***

### Running Unit Tests

To run unit tests, use the dbt test command with appropriate selectors:

```bash
# Run all unit tests
dbt test --select test_type:unit

# Run unit tests for a specific model
dbt test --select my_model,test_type:unit

# Run a specific unit test
dbt test --select my_specific_unit_test
```

***

### Testing Special Cases

#### Incremental Models

When testing incremental models, you can override the `is_incremental()` macro to test both full refresh and incremental scenarios:

```yaml
unit_tests:
  - name: test_incremental_full_refresh
    model: my_incremental_model
    overrides:
      macros:
        is_incremental: false
    # Test data here...

  - name: test_incremental_update
    model: my_incremental_model
    overrides:
      macros:
        is_incremental: true
    # Test data here including 'this' input...
```

For the incremental update test, you need to provide mock data for the existing table:

```yaml
given:
  - input: this  # Special reference to the current model
    rows:
      - {id: 1, value: 'existing', updated_at: '2023-01-01'}
      - {id: 2, value: 'existing', updated_at: '2023-01-01'}
```

#### Ephemeral Models

To test models that depend on ephemeral models, use SQL format for the input:

```yaml
unit_tests:
  - name: test_model_with_ephemeral_dependency
    model: my_model
    given:
      - input: ref('ephemeral_model')
        format: sql
        rows: |
          select 1 as id, 'test' as name
    # Expected output here...
```

#### Testing Macros

You can override macro implementations for testing:

```yaml
unit_tests:
  - name: test_with_custom_macro
    model: my_model
    overrides:
      macros:
        get_current_timestamp: return('2023-09-15')
    # Test data here...
```

***

### Limitations & Best Practices

dbt's unit testing framework has some limitations to be aware of:

{% hint style="warning" %}

#### Current Limitations

* Only supports SQL models (not Python models)
* Can only test models in your current project
* Doesn't support materialized views
* Doesn't support recursive SQL or introspective queries
* Requires all table names to be aliased in join logic {% endhint %}

#### Best Practices for Effective Unit Testing

| Practice                      | Description                                                      |
| ----------------------------- | ---------------------------------------------------------------- |
| Focus on logic, not functions | Test your business logic rather than built-in database functions |
| Use descriptive test names    | Clearly explain what each test is verifying                      |
| Test edge cases               | Include unusual scenarios your logic needs to handle             |
| Only mock what's needed       | Define only the columns relevant to your test                    |
| Run in development            | Use unit tests during development, not in production             |
| Use in CI/CD                  | Integrate unit tests into your CI/CD pipeline                    |

***

### Test-Driven Development with dbt

Unit tests enable a test-driven development (TDD) workflow for your data transformations:

1. **Write a test**: Define the expected behavior
2. **Run the test**: It should fail because the model doesn't exist or doesn't handle the case yet
3. **Implement the model**: Create or modify the model to pass the test
4. **Verify**: Run the test again to confirm it passes
5. **Refactor**: Clean up your implementation while keeping the tests passing

This approach helps ensure your models correctly implement business requirements from the start.

By adopting unit testing in your dbt workflow, you can catch issues earlier, document model behavior, and build more confidence in your data transformations.


# Working with Tags

Learn how to use tags in dbt™ to organize models, streamline execution, and improve workflow efficiency in analytics engineering.

Tags in dbt are powerful metadata labels that can be applied to various resources in your project. They enable flexible model selection, improved workflow management, and better project organization.

### What Are Tags?

Tags are simple text labels you can assign to models, sources, snapshots, and other dbt resources. They help you organize, categorize, and select specific subsets of your project for execution or documentation.

By leveraging tags, you can:

* **Group models logically** – Categorize models based on refresh schedule, function, or ownership
* **Control execution** – Run or exclude specific sets of models
* **Optimize CI/CD pipelines** – Target models for incremental builds and tests
* **Improve project maintainability** – Standardize workflows across teams

***

### How to Apply Tags

Tags can be applied in two primary ways: directly in model files or in your project configuration file.

#### 1. Defining Tags in a Model File

Tags can be assigned directly within SQL models using the `config()` function:

```sql
{{ config(
    tags=["finance", "daily_refresh"]
) }}

SELECT *
FROM {{ ref('stg_transactions') }}
```

This assigns both the "finance" and "daily\_refresh" tags to this specific model.

#### 2. Defining Tags in `dbt_project.yml`

Tags can also be applied at the project level, affecting entire folders or groups of models:

```yaml
models:
  my_project:
    +tags: "core"  # Assigns a single tag

    staging:
      +tags: ["staging", "raw_data"]  # Multiple tags

    marts:
      +tags:
        - "mart"
        - "business_logic"
```

***

### Tag Inheritance

Models inside a folder inherit the parent folder's tags unless overridden. This creates a hierarchical tagging system that is easy to maintain.

Example:

```
models/
  staging/
    customers.sql   → Tags: ["staging", "raw_data"]
    orders.sql      → Tags: ["staging", "raw_data"]
  marts/
    dim_customer.sql → Tags: ["mart", "business_logic"]
```

Individual models can add to inherited tags:

```sql
-- models/staging/special_orders.sql
{{ config(
    tags=["critical"]  # Adds to inherited ["staging", "raw_data"]
) }}

SELECT * FROM {{ ref('stg_orders') }}
```

This model would have tags: \["staging", "raw\_data", "critical"]

***

### Applying Tags to Different Resource Types

Tags can be applied to various dbt resource types:

#### Snapshots

```yaml
snapshots:
  my_project:
    +tags: ["historical_data"]
```

#### Seeds

```yaml
seeds:
  my_project:
    +tags: ["seed_data"]
```

#### Sources

```yaml
sources:
  - name: external_source
    tags: ['external']
```

***

### Using Tags for Selection

Once tags are defined, you can use them with dbt commands to select specific resources.

#### Selection Examples

| Command                                                       | Description                                                             |
| ------------------------------------------------------------- | ----------------------------------------------------------------------- |
| `dbt run --select tag:daily_refresh`                          | Run only models with the `daily_refresh` tag                            |
| `dbt run --select tag:daily_refresh tag:critical`             | Run models with either `daily_refresh` OR `critical` tags               |
| `dbt run --select tag:daily_refresh --exclude tag:deprecated` | Run models with `daily_refresh` but exclude those with `deprecated` tag |
| `dbt run --select staging,tag:finance`                        | Run all models tagged `finance` in the `staging` folder                 |
| `dbt run --select tag:critical+`                              | Run `critical` models and their downstream dependencies                 |

#### Tag Selection Patterns

| Selection Pattern     | What It Selects                              | Example                                     |
| --------------------- | -------------------------------------------- | ------------------------------------------- |
| `tag:name`            | All resources with this tag                  | `dbt run --select tag:nightly`              |
| `tag:name1 tag:name2` | Resources with either tag                    | `dbt run --select tag:nightly tag:critical` |
| `tag:name+`           | Tagged resources and downstream dependencies | `dbt run --select tag:base+`                |
| `+tag:name`           | Tagged resources and upstream dependencies   | `dbt run --select +tag:reporting`           |
| `--exclude tag:name`  | Everything except resources with this tag    | `dbt run --exclude tag:deprecated`          |

***

### Best Practices for Using Tags

#### Use Consistent Naming Conventions

Standardized naming improves clarity and prevents confusion.

```yaml
# Good
+tags: ["daily_refresh", "finance_data"]

# Avoid inconsistent casing or spacing
+tags: ["Daily", "Finance", "financial-data"]
```

#### Document Your Tagging Strategy

Clearly define tag meanings in your project's documentation.

```markdown
# Tag Definitions
- `daily_refresh`: Models refreshed daily.
- `finance_data`: Contains financial-related tables.
- `pii`: Includes personally identifiable information.
```

#### Use Granular Tags

Avoid broad, generic tags. Instead, use precise labels for better control.

```yaml
# Good
+tags: ["customer_metrics", "daily_refresh"]

# Too broad
+tags: ["metrics", "regular"]
```

#### Tag Models by Layer

Use tags to represent data modeling layers in your project.

```yaml
models:
  staging:
    +tags: ["bronze_layer"]
  intermediate:
    +tags: ["silver_layer"]
  marts:
    +tags: ["gold_layer"]
```

***

### Common Use Cases for Tags

#### Refresh Scheduling

Define tags based on refresh frequency for better execution control:

```yaml
models:
  my_project:
    hourly:
      +tags: ["hourly_refresh"]
    daily:
      +tags: ["daily_refresh"]
```

Then in your orchestration tool, schedule different runs:

```bash
# Morning run for daily models
dbt run --select tag:daily_refresh

# Every hour for hourly models
dbt run --select tag:hourly_refresh
```

#### Data Classification

Differentiate datasets based on sensitivity or access level:

```yaml
models:
  my_project:
    +tags: ["contains_pii"]
    public:
      +tags: ["public_data"]
```

This helps implement appropriate security controls and auditing.

#### Testing Strategy

Prioritize critical models in testing workflows:

```yaml
models:
  my_project:
    critical:
      +tags: ["critical_path", "requires_alert"]
```

Run critical tests more frequently:

```bash
dbt test --select tag:critical_path
```

***

### Troubleshooting Tag Issues

If dbt isn't selecting resources correctly based on tags, consider these troubleshooting steps:

#### Tag Inheritance Issues

Verify parent folder configurations in `dbt_project.yml`:

```bash
# List all models with their tags
dbt ls -m "my_project.staging.*" --output path
```

#### Selection Syntax Errors

Ensure tag names match exactly (case-sensitive):

```bash
# Try with exact case
dbt run --select tag:daily_refresh  # Not tag:Daily_Refresh
```

#### Using dbt ls to Validate Tags

Use the `dbt ls` command to check which models have specific tags:

```bash
# List all models with the "finance" tag
dbt ls --select tag:finance
```

By effectively implementing a tagging strategy, you can organize your dbt project more efficiently, streamline your workflows, and gain better control over how your transformations are executed.


# Managing Seeds

Learn how to use dbt seeds to manage static data. Discover best practices for creating, loading, configuring, and maintaining CSV files in your dbt project for seamless data integration.

Seeds are CSV files that live in your dbt project, typically containing static data that doesn't change frequently. They provide a convenient way to incorporate reference data, mapping tables, or small datasets directly into your dbt workflow.

***

### What Are Seeds?

Seeds are CSV files stored in your dbt project's `seeds/` directory that dbt loads into your data warehouse. They become part of your version-controlled project and are accessible in models via the `ref()` function, just like any other model.

Seeds work best for data that changes infrequently, is relatively small in size, benefits from version control, and doesn't require complex transformations before use.

***

### Common Use Cases for Seeds

#### Reference and Mapping Tables

Seeds are ideal for relatively static reference data such as country codes, currency conversion rates, product categorizations, department hierarchies, and calendar dimensions like fiscal years or holidays.

#### Configuration Tables

Use seeds as configuration tables for your transformations, including parameter tables, inclusion/exclusion lists, feature toggles, metric definitions, and valid value lists.

#### Test Data

Seeds provide a convenient way to manage test datasets, including sample customer profiles, test transactions, and expected results for complex calculations.

***

### Creating and Using Seeds

#### Adding a Seed to Your Project

To add a seed to your project:

1. Create a CSV file with headers in the first row
2. Save it in your project's `seeds/` directory
3. Ensure the file uses proper CSV formatting (commas as delimiters, proper escaping)

Example `seeds/country_codes.csv`:

```
country_code,country_name
US,United States
CA,Canada
MX,Mexico
UK,United Kingdom
DE,Germany
```

#### Loading Seeds into Your Warehouse

Run the seed command to load all seeds into your data warehouse:

```bash
dbt seed
```

Or load a specific seed:

```bash
dbt seed --select country_codes
```

#### Referencing Seeds in Models

Reference seeds in your models using the `ref()` function, just like you would with models:

```sql
select
  customer_id,
  customer_country_code,
  cc.country_name
from {{ ref('customers') }}
left join {{ ref('country_codes') }} as cc
  on customers.country_code = cc.country_code
```

***

### Configuring Seeds

#### Basic Configuration

Configure seeds in your `dbt_project.yml` file:

```yaml
seeds:
  my_project:
    country_codes:
      +column_types:
        country_code: varchar(2)
        country_name: varchar(100)
    +schema: reference_data
```

#### Advanced Configuration Options

| Configuration    | Purpose                                              | Example                                                                                                   |
| ---------------- | ---------------------------------------------------- | --------------------------------------------------------------------------------------------------------- |
| Column Types     | Explicitly define data types for columns             | <p><code>+column\_types:</code><br><code>amount: float</code><br><code>transaction\_date: date</code></p> |
| Schema Placement | Control which schema seeds are loaded into           | `+schema: reference_data`                                                                                 |
| Quote Policies   | Manage identifier quoting for columns and seed names | `+quote_columns: true`                                                                                    |

***

### Best Practices for Using Seeds

#### When to Use Seeds (and When Not To)

| Use seeds when                                  | Consider alternatives when                        |
| ----------------------------------------------- | ------------------------------------------------- |
| The dataset is small (generally <10,000 rows)   | The dataset is large                              |
| Data changes infrequently                       | Data requires frequent updates                    |
| Version control is beneficial                   | Complex transformations are needed                |
| The team needs to collaborate on reference data | External systems manage the authoritative version |

***

#### Performance Considerations

Keep seeds small to avoid version control and build time issues. Use appropriate column types to optimize storage and be mindful of dependencies on seeds throughout your project. For larger datasets or frequently changing data, consider alternatives to seeds.

{% hint style="info" %}
**Seed File Size Limits**

While there's no hard limit on seed file size, large CSV files (>1MB) can cause performance issues and version control challenges. For larger datasets, consider using external data loading tools, creating a source table instead, or splitting the data into multiple smaller seed files.
{% endhint %}

#### Maintenance Strategies

Document the purpose of each seed and establish clear ownership for seed maintenance. Create processes for updating seeds and consider automation for seeds that change more frequently but still warrant being included in your project.

***

### Seeds in Paradime

Paradime enhances the seed experience with visual CSV editing, data previews, version history tracking, and simplified update workflows.

#### Rainbow CSV Integration

Paradime integrates Rainbow CSV, which makes seed file management more efficient through:

* Color-coded columns for easier reading
* Visual alignment of CSV data
* Multi-cursor column editing
* Automatic CSV format validation

{% hint style="info" %}
For more details, see our [Rainbow CSV documentation](/app-help/documentation/integrations/code-ide/rainbow-csv).
{% endhint %}


# Environment Management

Learn how to set up and manage development, testing, and production environments in dbt projects.

Effective environment management is crucial for maintaining data quality and enabling collaborative development in dbt projects. Properly separating development, testing, and production environments allows teams to develop and test transformations without disrupting critical business processes.

***

### Why Environment Separation Matters

Imagine working in a single environment where both developers and business users access the same system. When an analytics engineer accidentally introduces a breaking change to an important model, it immediately impacts dashboards and reports that business leaders rely on. This scenario erodes trust in data and can lead to poor decision-making.

| Benefit                | Description                                                      |
| ---------------------- | ---------------------------------------------------------------- |
| Reduced risk           | Test changes without affecting production data or business users |
| Improved collaboration | Multiple team members can work simultaneously without conflicts  |
| Better quality control | Validate transformations before they reach decision-makers       |
| Deployment control     | Manage the release process deliberately with proper reviews      |

***

### Common Environment Approaches

There are several ways to implement environments in dbt, with no single consensus in the community. The approach you choose depends on factors like team size, administrative capacity, and cost considerations.

#### 1. One Database, Multiple Schemas

This is dbt's recommended approach: using different schemas within one data warehouse to separate environments, with each developer having their own development schema.

```yaml
# Example profiles.yml
my_project:
  target: dev
  outputs:
    dev:
      type: snowflake
      schema: dbt_jsmith  # Schema named after user
      # Other connection details
    prod:
      type: snowflake
      schema: analytics
      # Other connection details
```

This approach provides good isolation while minimizing administrative overhead.

#### 2. One Database per Environment

Another approach is to have separate databases for development and production, with individual schemas for each developer in the development database.

This keeps the production database clean and uncluttered, while providing clear separation between environments.

#### 3. One Database per Developer

In some cases, teams opt to give each developer a personal copy of the production database. Schemas remain identical across all databases, and developers can clone data or run models as needed.

This provides maximum isolation but can increase administrative burden and costs.

***

### Configuring Environments in dbt

dbt provides several mechanisms for managing environments:

#### Using profiles.yml

The primary way to configure environments is through your `profiles.yml` file:

```yaml
my_project:
  target: dev  # Default target
  outputs:
    dev:
      type: snowflake
      schema: dbt_jsmith
      # Other connection details
    test:
      type: snowflake
      schema: dbt_staging
      # Other connection details
    prod:
      type: snowflake
      schema: analytics
      # Other connection details
```

Switch between environments by specifying the target:

```bash
dbt run --target prod
```

#### Using Target Schemas

One powerful feature for environment management is dbt's target schema functionality, which appends the target name to your schema:

```yaml
# In dbt_project.yml
models:
  my_project:
    +schema: marketing  # Creates schema: marketing_dev, marketing_prod, etc.
```

This automatically creates separate schemas for each environment while maintaining consistent naming.

#### Environment-Specific Variables

Define variables that change based on the environment:

```yaml
# In dbt_project.yml
vars:
  dev:
    log_level: debug
    row_limit: 1000
  prod:
    log_level: info
    row_limit: null
```

Access these in your models:

```sql
SELECT *
FROM {{ ref('large_table') }}
{% if var('row_limit') %}
LIMIT {{ var('row_limit') }}
{% endif %}
```

***

### Key Considerations for Environment Management

When designing your environment strategy, consider these important factors:

<table><thead><tr><th width="171">Consideration</th><th>Details</th><th>Recommendations</th></tr></thead><tbody><tr><td>Administrative Burden</td><td>Managing multiple environments becomes complex with larger teams</td><td>• Use custom schemas with consistent naming<br><br>• Create self-service options for refreshing development data<br><br>• Automate environment provisioning and cleanup</td></tr><tr><td>Team Size &#x26; Collaboration</td><td>The number of developers affects your environment strategy</td><td>• Small teams (1-2 developers): Consider shared development environment<br><br>• Larger teams: Implement discrete spaces to avoid conflicts<br><br>• Enable read access to other development environments for collaboration</td></tr><tr><td>Cost Management</td><td>Multiple copies of production data increases compute and storage costs</td><td>• Use data filtering in development (e.g., <code>WHERE date >= CURRENT_DATE - 30</code>)<br><br>• Utilize database cloning features when available<br><br>• Evaluate if complete production copies are necessary<br><br>• Limit development compute resources</td></tr></tbody></table>

***

### Environment Management in Paradime

Paradime enhances environment management with several integrated features:

#### Project Environments

Paradime allows you to create and manage environments through its interface, with environment-specific settings, variables, access permissions, and logging.

#### Deployment Management

The platform streamlines deployment between environments through deployment packages, scheduled promotions, change tracking, and rollback capabilities.

{% hint style="info" %}
In Paradime, environments are designed to work with your chosen approach, whether you use schema-based separation or distinct databases. See [Connection documentation](/app-help/documentation/settings/connections) for details.
{% endhint %}

***

### Best Practices

Regardless of your chosen approach, follow these environment management best practices:

<table><thead><tr><th width="304">Environment</th><th>Best Practices</th></tr></thead><tbody><tr><td>Development</td><td>• Use consistent naming conventions<br>• Establish processes for regular environment cleanup<br>• Enable detailed logging for troubleshooting<br>• Set query timeouts to catch inefficient models early</td></tr><tr><td>Testing/Staging</td><td>• Maintain a near-exact replica of production<br>• Implement automated testing through CI/CD pipelines<br>• Regularly validate with production-like data volumes<br>• Perform performance testing before deployment</td></tr><tr><td>Production</td><td>• Implement change approval workflows<br>• Use appropriately sized compute resources<br>• Set up monitoring and alerting for performance and errors<br>• Maintain thorough documentation of the environment</td></tr></tbody></table>


# Variables and Parameters

Learn how to use variables and parameters in dbt to make your data   transformations configurable, flexible, and environment-aware.

Variables allow you to make your dbt project more dynamic and configurable by passing values at runtime or setting them in configuration files. They enable you to create flexible data transformations that can adapt to different environments, use cases, and scenarios.

### Understanding dbt Variables

Variables in dbt serve two primary purposes:

1. **Make code reusable** - Define values once and reference them throughout your project
2. **Enable flexibility** - Change behavior without modifying code

There are several ways to define and use variables in dbt:

* **Project variables** - Defined in `dbt_project.yml`
* **Command-line variables** - Passed at runtime
* **Environment variables** - Accessed via Jinja macros

***

### Defining Variables in dbt\_project.yml

The simplest way to define variables is in your `dbt_project.yml` file:

```yaml
vars:
  # Simple scalar values
  start_date: '2020-01-01'
  end_date: '2022-12-31'
  
  # Lists
  excluded_countries: ['test', 'demo', 'internal']
  
  # Dictionaries
  partner_sales_targets: {
    'tier1': 1000000,
    'tier2': 500000,
    'tier3': 100000
  }
  
  # Environment-specific variables
  dev:
    row_limit: 100
    debug_mode: true
  prod:
    row_limit: null
    debug_mode: false
```

These variables become available throughout your project via the `var()` function.

***

### Using Variables in Models

Once defined, you can reference variables in your models using the `var()` function:

```sql
-- models/reporting/monthly_sales.sql
SELECT
  date_trunc('month', order_date) as month,
  SUM(amount) as monthly_sales
FROM {{ ref('stg_orders') }}
WHERE 
  order_date >= '{{ var("start_date") }}'
  AND order_date <= '{{ var("end_date") }}'
  {% if var('row_limit') %}
  LIMIT {{ var('row_limit') }}
  {% endif %}
```

The `var()` function has two parameters:

1. The variable name
2. An optional default value that's used if the variable isn't defined

```sql
sqlCopy-- Using a default value
SELECT * FROM {{ ref('stg_users') }}
WHERE status = '{{ var("user_status", "active") }}'
```

{% hint style="info" %}
**Variable Behavior**

When you use the `var()` function:

* It will use the variable from `dbt_project.yml` if defined
* Command-line variables override values from `dbt_project.yml`
* If no variable is found and no default is specified, dbt will raise an error
* Environment-specific variables (`dev`, `prod`) are only used when running in that environment
  {% endhint %}

***

#### Passing Variables at Runtime

For maximum flexibility, pass variables at runtime using the `--vars` flag:

```bash
dbt run --vars '{"start_date": "2023-01-01", "end_date": "2023-03-31"}'
```

You can pass complex structures too:

```bash
dbt run --vars '{"regions": ["north", "south"], "include_test_data": false}'
```

Runtime variables override any variables defined in `dbt_project.yml`.

***

### Working with Environment Variables

You can access environment variables using the `env_var` Jinja function:

```sql
-- Configuring a model to use environment variables
{{ 
  config(
    schema=env_var('DBT_SCHEMA', 'analytics')
  ) 
}}

SELECT * FROM {{ ref('stg_orders') }}
```

This is particularly useful for sensitive information (like API keys) or values that vary by environment.

{% hint style="info" %}
**Security Note**

Never use `env_var()` for credentials that should remain secret. These values could be exposed in compiled SQL or logs. Instead, use your platform's secure environment variable handling for credentials.
{% endhint %}

***

### Advanced Variable Techniques

**Conditional Logic with Variables**

Variables allow you to implement conditional logic in your models:

```sql
{% if var('data_source') == 'api' %}
  SELECT * FROM {{ ref('stg_api_data') }}
{% else %}
  SELECT * FROM {{ ref('stg_warehouse_data') }}
{% endif %}
```

**Dynamic Filtering**

Create flexible filtering based on variable values:

```sql
SELECT
  *
FROM {{ ref('stg_transactions') }}
WHERE 1=1
  {% if var('filter_by_date', false) %}
  AND transaction_date BETWEEN '{{ var("start_date") }}' AND '{{ var("end_date") }}'
  {% endif %}
  
  {% if var('filter_by_country', false) %}
  AND country IN (
    {% for country in var('countries', []) %}
      '{{ country }}'{% if not loop.last %},{% endif %}
    {% endfor %}
  )
  {% endif %}
```

**Date/Time Variables**

A common pattern for incremental models is using variables for date ranges:

```sql
{% set run_date = var('run_date', modules.datetime.date.today().strftime('%Y-%m-%d')) %}

SELECT 
  *
FROM {{ source('events', 'daily_events') }}
WHERE 
  event_date = '{{ run_date }}'
```

***

#### Best Practices for Variables

| Practice                | Description                                                          |
| ----------------------- | -------------------------------------------------------------------- |
| Set meaningful defaults | Provide sensible default values to make your code more robust        |
| Use descriptive names   | Choose clear, explicit variable names that explain purpose           |
| Document variables      | Add comments in `dbt_project.yml` to explain each variable's purpose |
| Consistent formatting   | Maintain consistent casing and naming conventions                    |
| Avoid hardcoding        | Use variables instead of hardcoding values that might change         |

**Example: Well-Structured Variables**

```yaml
vars:
  # Analysis date range - used for filtering transaction data
  # Format: YYYY-MM-DD
  analysis_start_date: '2023-01-01'  # Inclusive
  analysis_end_date: '2023-12-31'    # Inclusive
  
  # Revenue recognition settings
  rev_rec_delay_days: 14             # Days to delay revenue recognition
  include_refunds: false             # Set to true to include refunded transactions
  
  # Environment-specific settings
  dev:
    debug_mode: true                 # Enables additional logging
    data_sample_pct: 10              # Only process 10% of data in dev
  prod:
    debug_mode: false
    data_sample_pct: 100             # Process all data in prod
```

***

### Common Use Cases

**Environment-Specific Configuration**

Define different behavior based on your deployment environment:

```yaml
# dbt_project.yml
vars:
  dev:
    schema_prefix: 'dev_'
    row_limit: 1000
  prod:
    schema_prefix: ''
    row_limit: null
```

```sql
-- models/model.sql
{{ 
  config(
    schema=var('schema_prefix', 'dev_') ~ 'marketing'
  ) 
}}

SELECT * FROM {{ ref('stg_data') }}
{% if var('row_limit') %}
LIMIT {{ var('row_limit') }}
{% endif %}
```

**Parameterized Reporting**

Create reports with customizable parameters:

```sql
-- models/daily_sales_report.sql
{% set date_column = var('date_column', 'order_date') %}
{% set granularity = var('granularity', 'day') %}

SELECT
  DATE_TRUNC('{{ granularity }}', {{ date_column }}) as period,
  SUM(amount) as sales
FROM {{ ref('fct_orders') }}
GROUP BY 1
ORDER BY 1
```

Then run with different settings:

```bash
dbt run --select daily_sales_report --vars '{"granularity": "month", "date_column": "shipped_date"}'
```

By effectively using variables in your dbt project, you create more flexible, maintainable, and reusable data transformations that can easily adapt to different needs and environments without code changes.


# Macros

Learn how to use dbt™ macros to create reusable SQL logic and build    modular transformations. This guide covers creating macros, implementing   advanced techniques, and leveraging dbt packages to ex

Macros are powerful features that allow you to create reusable code patterns and implement dynamic SQL generation in your dbt projects. This guide will help you understand how to use macros to make your dbt projects more maintainable, consistent, and flexible.

### What Are Macros?

Macros are reusable pieces of code that let you eliminate repetition, create project-wide standards, and abstract complex logic. Think of macros as functions in traditional programming languages that can be called from other macros, models, or schema files.

Macros enable you to:

* Abstract complex SQL logic into reusable functions
* Create project-wide standards for common calculations
* Implement conditional logic in your SQL code
* Generate SQL dynamically based on parameters

***

### Creating Your First Macro

Macros are defined in `.sql` files within the `macros` directory of your dbt project. A basic macro follows this structure:

```sql
{% macro macro_name(parameter1, parameter2, ...) %}
    
    -- SQL code and Jinja logic goes here
    
    {% if parameter1 > 0 %}
        SELECT {{ parameter1 }} + {{ parameter2 }}
    {% else %}
        SELECT {{ parameter2 }}
    {% endif %}
    
{% endmacro %}
```

**Example: Creating a Date Dimension Macro**

Here's a practical example of a macro that generates a date dimension table:

```sql
{% macro generate_date_dimension(start_date, end_date) %}

    WITH date_spine AS (
        {{ dbt_utils.date_spine(
            datepart="day",
            start_date="cast('" ~ start_date ~ "' as date)",
            end_date="cast('" ~ end_date ~ "' as date)"
        ) }}
    ),
    
    dates AS (
        SELECT
            cast(date_day as date) as date_day,
            extract(year from date_day) as year,
            extract(month from date_day) as month,
            extract(day from date_day) as day_of_month,
            extract(dayofweek from date_day) as day_of_week,
            extract(quarter from date_day) as quarter
        FROM date_spine
    )
    
    SELECT * FROM dates

{% endmacro %}
```

This macro leverages the `date_spine` utility from dbt\_utils to create a complete date dimension table with various date attributes.

**Using Macros in Your Models**

To use a macro in a model, you simply call it using the Jinja templating syntax:

```sql
-- models/dim_date.sql
{{
    config(
        materialized='table'
    )
}}

{{ generate_date_dimension('2020-01-01', '2025-12-31') }}
```

When dbt runs this model, it will replace the macro call with the SQL generated by the macro, creating a date dimension table for the specified date range.

***

### Advanced Macro Techniques

**Macro Organization**

For larger projects, organizing macros in subdirectories helps maintain a clean structure:

```
macros/
  ├── date_utils/
  │   ├── generate_date_dimension.sql
  │   └── fiscal_year_dates.sql
  ├── string_utils/
  │   ├── clean_string.sql
  │   └── standardize_phone.sql
  └── metrics/
      ├── calculate_revenue.sql
      └── customer_lifetime_value.sql
```

**Using Control Structures**

Macros support Jinja's control structures for advanced logic:

```sql
{% macro dynamic_pivot(table_name, group_by_columns, pivot_column, value_column) %}

    {% set group_by_str = group_by_columns | join(', ') %}
    
    {% set query %}
        SELECT DISTINCT {{ pivot_column }} 
        FROM {{ table_name }}
        ORDER BY 1
    {% endset %}
    
    {% set results = run_query(query) %}
    
    {% if execute %}
        {% set pivot_values = results.columns[0].values() %}
    {% else %}
        {% set pivot_values = [] %}
    {% endif %}
    
    SELECT
        {{ group_by_str }},
        {% for value in pivot_values %}
            SUM(CASE WHEN {{ pivot_column }} = '{{ value }}' THEN {{ value_column }} ELSE 0 END) AS "{{ value }}"
            {% if not loop.last %},{% endif %}
        {% endfor %}
    FROM {{ table_name }}
    GROUP BY {{ group_by_str }}
    
{% endmacro %}
```

This advanced macro dynamically creates a pivot table based on values found in your data at runtime.

{% hint style="info" %}
**Using Macros Effectively**

* **Keep macros focused** – Each macro should do one thing well
* **Document your macros** – Add comments explaining parameters and usage
* **Use Jinja's `execute` flag** – The code within `{% if execute %}` only runs during compilation, not during preview or rendering
* **Test macros thoroughly** – Create models specifically for testing macro functionality
* **Use default parameters** – Make macros flexible while providing sensible defaults
  {% endhint %}

***

### Working with dbt Packages

dbt packages let you leverage pre-built macros created by the community. They're an excellent way to avoid reinventing the wheel.

**Installing Packages**

To use packages, define them in a `packages.yml` file in your project root:

```yaml
packages:
  - package: dbt-labs/dbt_utils
    version: 1.1.1
  - package: calogica/dbt_expectations
    version: 0.8.5
```

Then install the packages using:

```bash
dbt deps
```

**Popular Packages for Macros**

| Package         | Purpose                         | Key Features                                               |
| --------------- | ------------------------------- | ---------------------------------------------------------- |
| **dbt-utils**   | General utility macros          | String operations, date handling, cross-database functions |
| **dbt-date**    | Date and calendar functionality | Date spines, fiscal periods, date utilities                |
| **dbt-ml**      | Machine learning functionality  | Feature engineering, model scoring                         |
| **dbt-codegen** | Code generation tools           | Auto-generate models, sources, base models                 |

**Using Package Functions**

Once installed, you can use package functions in your models:

```sql
-- Example using dbt_utils generate_surrogate_key
SELECT
    {{ dbt_utils.generate_surrogate_key(['customer_id', 'order_date']) }} as order_key,
    customer_id,
    order_date,
    amount
FROM {{ ref('stg_orders') }}
```

***

### Real-World Examples

**Financial Calculations**

```sql
{% macro calculate_margin(revenue, cost) %}
    ({{ revenue }} - {{ cost }}) / NULLIF({{ revenue }}, 0)
{% endmacro %}
```

**Dynamic Table Generation**

```sql
{% macro generate_surrogate_key(field_list) %}
    {% set fields = [] %}
    {% for field in field_list %}
        {% do fields.append("coalesce(cast(" ~ field ~ " as string), '')") %}
    {% endfor %}
    md5({{ fields|join(" || '-' || ") }})
{% endmacro %}
```

**Environment-Based Configuration**

```sql
{% macro get_schema_prefix() %}
    {% if target.name == 'prod' %}
        prod_
    {% elif target.name == 'dev' %}
        dev_{{ target.schema }}_
    {% else %}
        {{ target.schema }}_
    {% endif %}
{% endmacro %}
```

***

#### Best Practices for Macros

| Best Practice              | Description                                                                                               |
| -------------------------- | --------------------------------------------------------------------------------------------------------- |
| **Keep Macros Focused**    | Each macro should do one thing well. Avoid overly complex macros with many responsibilities.              |
| **Document Your Macros**   | Add comments explaining purpose, parameters, and return values. Include examples of how to use the macro. |
| **Test Your Macros**       | Create models specifically for testing macro functionality. Use assertions to verify macro outputs.       |
| **Handle Edge Cases**      | Ensure your macros handle null values appropriately. Account for empty tables and edge conditions.        |
| **Use Default Parameters** | Make macros flexible with sensible defaults. Allow override of defaults when needed.                      |
| **Leverage Return Values** | Use `return()` to pass values back from macros. Chain macros together for complex operations.             |

{% hint style="info" %}
**Pro Tip: Debugging Macros**

When troubleshooting macros:

1. Use `dbt compile` to see the generated SQL without running it
2. Check the compiled SQL in the `target/compiled/` directory
3. Add `{{ log("Debug message") }}` within macros for debugging
4. Use `{% if execute %}` to handle compile-time vs. run-time logic
   {% endhint %}

By mastering macros, you can create more maintainable, consistent, and flexible data transformations throughout your dbt project.


# Custom Tests

Learn how to extend dbt™'s testing capabilities with custom tests. This guide covers   creating singular and generic tests, implementing parameterized tests, and integrating   with dbt packages for ad

While dbt comes with built-in tests, custom tests allow you to implement specific data quality checks tailored to your business logic. This guide will help you understand how to create and use custom tests to ensure the quality and reliability of your data transformations.

#### Types of Custom Tests

dbt supports two main types of custom tests:

| Test Type      | Description                                              | Where Defined       | Scope             |
| -------------- | -------------------------------------------------------- | ------------------- | ----------------- |
| Singular Tests | SQL files returning failing records                      | `tests/` directory  | Specific use case |
| Generic Tests  | Reusable test definitions applicable to different models | `macros/` directory | Reusable          |

***

### Singular Tests

Singular tests are SQL queries that should return zero rows when the test passes:

```sql
-- tests/assert_total_payment_amount_matches_order_amount.sql

SELECT
    order_id,
    order_amount,
    payment_amount,
    ABS(order_amount - payment_amount) as amount_diff
FROM {{ ref('orders') }} o
LEFT JOIN {{ ref('payments') }} p USING (order_id)
WHERE ABS(order_amount - payment_amount) > 0.01
```

This test identifies orders where the payment amount doesn't match the order amount within a small tolerance.

To run singular tests:

```bash
dbt test --select assert_total_payment_amount_matches_order_amount
```

**Organizing Singular Tests**

For larger projects, you might want to organize singular tests by domain or purpose:

```
tests/
  ├── finance/
  │   ├── assert_total_payment_amount_matches_order_amount.sql
  │   └── assert_refund_amount_less_than_order_amount.sql
  └── marketing/
      ├── assert_campaign_spend_within_budget.sql
      └── assert_conversion_rates_above_threshold.sql
```

***

### Creating Generic Custom Tests

Generic tests are more powerful because they can be applied to different models throughout your project. They are defined as macros with a special syntax:

```sql
{% test test_name(model, column_name, condition_parameter) %}

    -- The test query should return failing records
    SELECT
        {{ column_name }}
    FROM {{ model }}
    WHERE -- Test condition that should return 0 rows when passing
        {{ column_name }} NOT {{ condition_parameter }}

{% endtest %}
```

**Example: Positive Values Test**

Here's a simple custom test that checks if values in a column are positive:

```sql
{% test is_positive(model, column_name) %}

    SELECT
        {{ column_name }}
    FROM {{ model }}
    WHERE {{ column_name }} <= 0
    
{% endtest %}
```

**Using Custom Tests in YAML Files**

Once defined, you can reference these custom tests in your schema YAML files just like built-in tests:

```yaml
models:
  - name: orders
    columns:
      - name: order_amount
        tests:
          - is_positive
```

**Parameterizing Custom Tests**

You can make your custom tests more flexible by adding parameters:

```sql
{% test value_within_range(model, column_name, min_value, max_value) %}

    SELECT
        {{ column_name }}
    FROM {{ model }}
    WHERE {{ column_name }} < {{ min_value }} OR {{ column_name }} > {{ max_value }}
    
{% endtest %}
```

In your YAML file:

```yaml
models:
  - name: orders
    columns:
      - name: order_amount
        tests:
          - value_within_range:
              min_value: 0
              max_value: 10000
```

***

### Combining Macros and Tests

You can use macros within your custom tests to create powerful, reusable testing frameworks:

```sql
{% macro get_valid_status_values() %}
    {% set valid_statuses = ['pending', 'shipped', 'delivered', 'cancelled'] %}
    {{ return(valid_statuses) }}
{% endmacro %}

{% test valid_status(model, column_name) %}

    {% set valid_values = get_valid_status_values() %}
    
    SELECT
        {{ column_name }}
    FROM {{ model }}
    WHERE {{ column_name }} NOT IN (
        {% for value in valid_values %}
            '{{ value }}'{% if not loop.last %},{% endif %}
        {% endfor %}
    )
    
{% endtest %}
```

This approach lets you centralize business rules (like valid status values) and reuse them across tests.

***

### Advanced Test Configurations

You can configure how tests behave using additional properties:

```sql
{% test custom_complex_test(model, column_name) %}

    {{ config(
        severity = 'warn',
        store_failures = true,
        limit = 100
    ) }}

    -- Test query here
    
{% endtest %}
```

| Configuration       | Description                                     | Example Use Case                                  |
| ------------------- | ----------------------------------------------- | ------------------------------------------------- |
| **severity**        | Can be 'error' (default) or 'warn'              | For tests that shouldn't block production runs    |
| **store\_failures** | When true, stores test failures in a table      | For troubleshooting or monitoring over time       |
| **limit**           | Maximum number of failing records to return     | For large tables where full results aren't needed |
| **where**           | Apply additional filtering to the test query    | For focusing tests on specific data subsets       |
| **enabled**         | Boolean that can conditionally disable the test | For environment-specific test configuration       |

***

### Testing Data Quality with Packages

Popular packages like dbt-expectations extend dbt's testing capabilities with advanced data validation:

```yaml
packages:
  - package: calogica/dbt_expectations
    version: 0.8.5
```

**Example: Advanced Data Validation**

Using dbt-expectations to implement sophisticated data quality checks:

```yaml
models:
  - name: customer_orders
    tests:
      - dbt_expectations.expect_table_row_count_to_be_between:
          min_value: 1
          max_value: 1000000
    columns:
      - name: order_amount
        tests:
          - dbt_expectations.expect_column_values_to_be_between:
              min_value: 0
              max_value: 50000
              severity: warn
      - name: customer_email
        tests:
          - dbt_expectations.expect_column_values_to_match_regex:
              regex: '^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$'
```

**Popular Testing Packages**

| Package              | Purpose                                           | Key Features                                       |
| -------------------- | ------------------------------------------------- | -------------------------------------------------- |
| **dbt-expectations** | Data quality tests inspired by Great Expectations | Advanced schema validation, statistical tests      |
| **dbt-audit-helper** | Compare query results between models              | Model comparison, reconciliation tests             |
| **dbt-utils**        | Common test utilities                             | Equal rowcounts, relationships, cardinality checks |
| **elementary**       | Anomaly detection and data validation             | Historical test comparisons, metrics monitoring    |

***

### Real-World Custom Test Examples

**Date Range Validation**

```sql
{% test date_between_project_dates(model, column_name) %}
    SELECT
        {{ column_name }}
    FROM {{ model }}
    WHERE {{ column_name }} NOT BETWEEN 
        (SELECT min_valid_date FROM {{ ref('project_date_settings') }})
        AND
        (SELECT max_valid_date FROM {{ ref('project_date_settings') }})
{% endtest %}
```

**Numeric Distribution Test**

```sql
{% test standard_deviation_within_range(model, column_name, min_stddev, max_stddev) %}
    WITH stats AS (
        SELECT 
            STDDEV({{ column_name }}) AS std_dev
        FROM {{ model }}
        WHERE {{ column_name }} IS NOT NULL
    )
    
    SELECT *
    FROM stats
    WHERE std_dev < {{ min_stddev }} OR std_dev > {{ max_stddev }}
{% endtest %}
```

**Referential Integrity with Exceptions**

```sql
{% test foreign_key_with_exceptions(model, column_name, to, field, exceptions) %}
    
    {% set exceptions_list = [] %}
    {% for exception in exceptions %}
        {% do exceptions_list.append("'" ~ exception ~ "'") %}
    {% endfor %}
    
    SELECT
        {{ column_name }}
    FROM {{ model }}
    WHERE {{ column_name }} NOT IN (
        SELECT {{ field }}
        FROM {{ to }}
    )
    AND {{ column_name }} NOT IN ({{ exceptions_list | join(', ') }})
    
{% endtest %}
```

***

### Best Practices for Custom Tests

<table><thead><tr><th width="261.0625">Best Practice</th><th>Description</th></tr></thead><tbody><tr><td><strong>Test Critical Data First</strong></td><td>Focus on testing key business metrics and join keys. Identify high-risk areas for data quality issues.</td></tr><tr><td><strong>Make Tests Descriptive</strong></td><td>Name tests clearly to indicate what they verify. Add comments explaining the purpose and expectations.</td></tr><tr><td><strong>Balance Coverage and Performance</strong></td><td>Consider the runtime impact of extensive testing. Use selective testing for large tables.</td></tr><tr><td><strong>Group Related Tests</strong></td><td>Organize tests that verify related business rules. Use consistent naming conventions.</td></tr><tr><td><strong>Handle Edge Cases</strong></td><td>Test boundary conditions. Consider null handling and empty tables.</td></tr><tr><td><strong>Monitor Test Results</strong></td><td>Track test failures over time. Establish alerting for critical test failures.</td></tr></tbody></table>

{% hint style="info" %}
**Pro Tip: Troubleshooting Failed Tests**

When tests fail, use these strategies to diagnose issues:

1. Check error messages in the dbt logs
2. Examine a sample of failing records with `store_failures: true`
3. Verify test logic by inspecting the compiled SQL in `target/compiled/`
4. Use `--vars` to test with different parameters
   {% endhint %}

By implementing custom tests, you can ensure your transformations meet business requirements and maintain high data quality standards throughout your dbt project.


# Hooks & Operational Tasks

Learn how to use dbt™ hooks to automate operational tasks like granting permissions,   managing metadata, and implementing custom logic before or after model runs.

Hooks in dbt allow you to execute SQL statements or custom logic at specific points in your dbt workflow. They're powerful tools for automating operational tasks, managing permissions, and enhancing your data pipeline with custom functionality.

### What Are Hooks?

Hooks are snippets of SQL that run at predefined moments during the dbt execution process. They help automate operational tasks like:

* Granting permissions on objects
* Setting table properties or attributes
* Adding comments or metadata
* Creating indexes or optimizing data structures
* Executing database-specific operations

There are four main types of hooks in dbt:

| Hook Type      | Execution Timing                                      | Scope               | Use Cases                              |
| -------------- | ----------------------------------------------------- | ------------------- | -------------------------------------- |
| `pre-hook`     | Before a model, seed, or snapshot is built            | Model-specific      | Data validation, temporary table setup |
| `post-hook`    | After a model, seed, or snapshot is built             | Model-specific      | Adding indexes, granting permissions   |
| `on-run-start` | At the start of dbt commands (run, build, test, etc.) | Project/folder wide | Session configuration, global setup    |
| `on-run-end`   | At the end of dbt commands (run, build, test, etc.)   | Project/folder wide | Cleanup, notifications, audit logging  |

***

### Using Model-Level Hooks

**Pre-hooks and Post-hooks**

You can define pre-hooks and post-hooks directly in your model SQL files using the `config()` function:

```sql
-- models/customers.sql
{{ 
  config(
    materialized='table',
    pre_hook="DELETE FROM {{ this }} WHERE customer_id < 0",
    post_hook="GRANT SELECT ON {{ this }} TO ROLE analyst"
  ) 
}}

SELECT * FROM {{ ref('stg_customers') }}
```

In this example:

* The pre-hook deletes invalid records before the model builds
* The post-hook grants select permissions after the model is created

**Multiple Hooks**

You can specify multiple hooks as a list:

```sql
{{ 
  config(
    post_hook=[
      "GRANT SELECT ON {{ this }} TO ROLE analyst",
      "ANALYZE TABLE {{ this }}",
      "ALTER TABLE {{ this }} ADD COMMENT '{{ doc('customers_table_description') }}'"
    ]
  ) 
}}
```

***

### Project and Folder-Level Hooks

You can configure hooks that apply to groups of models in your `dbt_project.yml` file:

```yaml
models:
  my_project:
    # Hooks for all models
    +post_hook: "GRANT SELECT ON {{ this }} TO ROLE reporter"
    
    # Hooks for specific folders
    marts:
      +post_hook: "ANALYZE TABLE {{ this }}"
    
    staging:
      +pre_hook: "SET query_tag = 'staging_models'"
```

**On-Run-Start and On-Run-End Hooks**

These hooks run once at the beginning or end of your dbt run:

```yaml
on-run-start:
  - "SET timezone = 'America/Los_Angeles'"
  - "SET query_tag = 'dbt_run_{{ run_started_at.strftime('%Y%m%d_%H%M%S') }}'"

on-run-end:
  - "CALL audit.log_dbt_run('{{ run_started_at }}', '{{ invocation_id }}')"
```

{% hint style="info" %}
**Hook Context Variables**

In hooks, you can access several useful context variables:

* `{{ this }}` - The relation being built (table/view)
* `{{ target }}` - Information about the current target database
* `{{ run_started_at }}` - Timestamp when the run started
* `{{ invocation_id }}` - Unique ID for the current dbt run
  {% endhint %}

***

### Using Macros in Hooks

You can make your hooks more reusable by calling macros:

```sql
-- models/customers.sql
{{ 
  config(
    post_hook="{{ grant_select(this, 'analyst') }}"
  ) 
}}

SELECT * FROM {{ ref('stg_customers') }}
```

With a corresponding macro defined:

```sql
-- macros/grant_select.sql
{% macro grant_select(relation, role) %}
    grant select on {{ relation }} to role {{ role }};
{% endmacro %}
```

This approach allows you to centralize and reuse your hook logic across multiple models.

You can also configure hooks with macros in your YAML files:

```yaml
# models/schema.yml
models:
  - name: customers
    config:
      post_hook: "{{ grant_select(this, 'analyst') }}"
      
# dbt_project.yml
models:
  my_project:
    +post_hook: "{{ grant_select(this, 'reporter') }}"
```

***

### Common Hook Use Cases

**Permission Management**

A common use for hooks is automating permission grants:

```sql
-- Granting multiple permissions
{{ 
  config(
    post_hook=[
      "GRANT SELECT ON {{ this }} TO ROLE analyst",
      "GRANT SELECT ON {{ this }} TO ROLE reporter"
    ]
  ) 
}}
```

**Database-Specific Operations**

Perform operations specific to your data warehouse:

```sql
-- Snowflake example
{{ 
  config(
    post_hook="ALTER TABLE {{ this }} SET DATA_RETENTION_TIME_IN_DAYS = 90"
  ) 
}}

-- Redshift example
{{ 
  config(
    post_hook="VACUUM {{ this }}"
  ) 
}}
```

**Environment-Specific Actions**

Apply different hooks based on your deployment environment:

```sql
{{ 
  config(
    post_hook=
      {% if target.name == 'prod' %}
        "GRANT SELECT ON {{ this }} TO ROLE business_users"
      {% else %}
        "GRANT SELECT ON {{ this }} TO ROLE dbt_developers"
      {% endif %}
  ) 
}}
```

***

### Operations with run-operation

Operations are a way to execute standalone macros using the `run-operation` command. This is useful for administrative tasks that you want to run on demand, rather than as part of a model build.

**Creating an Operation Macro**

To create an operation, define a macro that performs the desired actions:

```sql
-- macros/grant_select.sql
{% macro grant_select(role) %}
    {% set sql %}
        grant usage on schema {{ target.schema }} to role {{ role }};
        grant select on all tables in schema {{ target.schema }} to role {{ role }};
        grant select on all views in schema {{ target.schema }} to role {{ role }};
    {% endset %}

    {% do run_query(sql) %}
    {% do log("Privileges granted", info=True) %}
{% endmacro %}
```

Note two important points:

1. The SQL is defined within a `{% set sql %}` block
2. The `run_query()` function is used to actually execute the SQL

**Running an Operation**

To run this operation from the command line:

```bash
dbt run-operation grant_select --args '{role: reporter}'
```

This would grant select privileges on all tables in your schema to the 'reporter' role.

{% hint style="info" %}
**Key Difference Between Hooks and Operations**

* **Hooks** are automatically executed at specific times during dbt runs
* **Operations** are explicitly run on-demand using the `run-operation` command
* With operations, you must use `run_query()` or a statement block to execute the SQL
  {% endhint %}

***

### Operation Examples

**Refreshing a Snowflake Pipe**

```sql
{% macro refresh_pipe(pipe_name) %}
    {% set sql %}
        ALTER PIPE {{ pipe_name }} REFRESH;
    {% endset %}
    
    {% do run_query(sql) %}
    {% do log("Pipe refreshed: " ~ pipe_name, info=True) %}
{% endmacro %}
```

**Creating Multiple Objects**

```sql
{% macro setup_monitoring() %}
    {% set sql %}
        CREATE SCHEMA IF NOT EXISTS {{ target.schema }}_monitor;
        
        CREATE TABLE IF NOT EXISTS {{ target.schema }}_monitor.audit_log (
            event_time TIMESTAMP_NTZ DEFAULT CURRENT_TIMESTAMP(),
            event_type VARCHAR(100),
            model_name VARCHAR(100),
            duration_seconds FLOAT
        );
    {% endset %}
    
    {% do run_query(sql) %}
    {% do log("Monitoring setup complete", info=True) %}
{% endmacro %}
```

**Passing Complex Arguments**

You can pass complex arguments to operations:

```bash
dbt run-operation create_test_data --args '{"schema": "analytics", "tables": ["customers", "orders"], "row_count": 1000}'
```

With a corresponding macro:

```sql
{% macro create_test_data(schema, tables, row_count) %}
    {% for table in tables %}
        {% set sql %}
            INSERT INTO {{ schema }}.{{ table }} /* Generate test data SQL here */
        {% endset %}
        {% do run_query(sql) %}
        {% do log("Generated " ~ row_count ~ " rows for " ~ schema ~ "." ~ table, info=True) %}
    {% endfor %}
{% endmacro %}
```

***

### Best Practices

#### Hooks

| Best Practice                          | Description                                                           |
| -------------------------------------- | --------------------------------------------------------------------- |
| **Use macros for repeated hook logic** | Create reusable macro functions instead of duplicating hook SQL.      |
| **Keep hooks focused**                 | Each hook should do one thing well.                                   |
| **Consider hook execution order**      | Remember that project-level hooks run before/after model-level hooks. |
| **Be careful with transactions**       | Understand your database's transaction behavior with hooks.           |
| **Test hooks in development**          | Verify hook behavior before deploying to production.                  |

#### Operations

| Best Practice                                    | Description                                               |
| ------------------------------------------------ | --------------------------------------------------------- |
| **Always use `run_query()` or statement blocks** | Operations must explicitly execute SQL.                   |
| **Add logging**                                  | Use `log()` to provide feedback about operation progress. |
| **Handle errors gracefully**                     | Consider `try/except` patterns for complex operations.    |
| **Document operation parameters**                | Make it clear what arguments your operations accept.      |
| **Use for administrative tasks**                 | Operations are perfect for one-off maintenance tasks.     |

{% hint style="info" %}
**When to Use Hooks vs. Operations**

Use **hooks** when you need to:

* Execute SQL automatically at specific points in your dbt workflow
* Apply consistent actions across multiple models
* Implement pre/post processing that's tightly coupled to models

Use **operations** when you need to:

* Run administrative tasks on-demand
* Perform one-off database maintenance
* Execute complex logic that doesn't fit into the model build process
* Create setup/teardown scripts for your environment
  {% endhint %}

By mastering hooks and operations, you can significantly extend dbt's capabilities and automate many aspects of database administration and data pipeline management.


# Packages

Learn how to extend dbt™'s functionality with packages. This guide covers installing, using, and creating packages to add reusable functionality like macros, models, and tests to your dbt project

dbt packages allow you to import pre-built models, macros, and tests into your project, helping you solve common data modeling challenges without reinventing the wheel. This guide explains how to use, install, and create dbt packages.

### What Are dbt Packages?

dbt packages are essentially standalone dbt projects that can be imported into your project. They contain reusable models, macros, tests, and other resources that extend dbt's functionality and help solve common data modeling challenges.

Packages enable you to:

* Leverage community-contributed solutions
* Standardize transformations across projects
* Import specialized functionality for specific data sources
* Apply consistent testing patterns
* Avoid reinventing solutions for common problems

***

### Adding Packages to Your Project

Using packages in your dbt project is a simple three-step process:

1. Create a `packages.yml` file in your project root (next to your `dbt_project.yml`)
2. Define the packages you want to use
3. Run `dbt deps` to install the packages

**Basic Package Configuration**

```yaml
# packages.yml
packages:
  - package: dbt-labs/dbt_utils
    version: 1.1.1
  
  - package: calogica/dbt_expectations
    version: 0.8.5
```

When you run `dbt deps`, dbt will install these packages into a `dbt_packages/` directory in your project. By default, this directory is ignored by git to avoid duplicating code.

***

### Package Installation Methods

dbt supports several methods for specifying package sources, depending on where your package is stored.

**Hub Packages (Recommended)**

The simplest way to install packages is from the dbt Hub:

```yaml
packages:
  - package: dbt-labs/snowplow
    version: 0.7.3
```

You can also specify version ranges using semantic versioning:

```yaml
packages:
  - package: dbt-labs/snowplow
    version: [">=0.7.0", "<0.8.0"]
```

This approach is recommended because the Hub can handle duplicate dependencies automatically.

**Git Packages**

For packages stored in Git repositories:

```yaml
packages:
  - git: "https://github.com/dbt-labs/dbt-utils.git"
    revision: 0.9.2
```

The `revision` parameter can be:

* A branch name
* A tag name
* A specific commit (40-character hash)

**Local Packages**

For packages on your local filesystem:

```yaml
packages:
  - local: relative/path/to/package
```

This is useful for testing package changes or working with monorepos.

{% hint style="info" %}
**Package Versioning Best Practices**

* Always pin package versions in production projects
* Use semantic versioning ranges for minor updates
* Test package updates thoroughly before deploying to production
* Beginning with dbt v1.7, running `dbt deps` automatically pins packages by creating a `package-lock.yml` file
  {% endhint %}

***

### Using Package Functionality

Once installed, you can use the resources from packages in your project.

**Using Package Macros**

Call macros from the package in your models:

```sql
-- Using dbt_utils.generate_surrogate_key
SELECT
    {{ dbt_utils.generate_surrogate_key(['customer_id', 'order_date']) }} as order_sk,
    customer_id,
    order_date,
    amount
FROM {{ ref('stg_orders') }}
```

**Using Package Tests**

Apply tests provided by packages in your schema files:

```yaml
models:
  - name: customers
    columns:
      - name: email
        tests:
          - dbt_expectations.expect_column_values_to_match_regex:
              regex: '^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$'
```

**Referencing Package Models**

Reference models from packages using the standard `ref` function:

```sql
-- Reference a model from a package
SELECT * FROM {{ ref('snowplow', 'snowplow_page_views') }}
```

When referencing models from packages, you can include the package name as the first argument to `ref`.

***

### Configuring Packages

Many packages allow you to configure their behavior using variables in your `dbt_project.yml` file:

```yaml
# dbt_project.yml

vars:
  # Configure the snowplow package
  snowplow:
    'snowplow:timezone': 'America/New_York'
    'snowplow:page_ping_frequency': 10
    'snowplow:events': "{{ ref('sp_base_events') }}"

# Override package configurations
models:
  snowplow:
    +schema: snowplow_models
```

You can also override materializations, schemas, or other configurations defined in the package.

***

### Popular dbt Packages

Here are some widely-used packages that can enhance your dbt projects:

| Package              | Purpose                    | Key Features                                              |
| -------------------- | -------------------------- | --------------------------------------------------------- |
| **dbt-utils**        | General utilities          | Cross-database macros, SQL helpers, schema tests          |
| **dbt-expectations** | Data quality testing       | Advanced testing functions inspired by Great Expectations |
| **dbt-date**         | Date/time functionality    | Date spine generation, fiscal periods, holiday calendars  |
| **dbt-audit-helper** | Auditing and comparison    | Model comparison, reconciliation helpers                  |
| **codegen**          | Code generation            | Auto-generate source definitions and base models          |
| **dbt-meta-testing** | Document and test coverage | Test your documentation and test coverage                 |

***

### Working with Private Packages

For organizations with internal packages, dbt supports several methods for authentication.

#### **Private Hub Packages**

You can use private packages with the proper authentication:

```yaml
packages:
  - private: dbt-labs/internal-package
    provider: "github"  # Specify if you have multiple git providers configured
```

**Git Token Method**

For HTTPS authentication with a token:

```yaml
packages:
  - git: "https://{{env_var('GIT_CREDENTIAL')}}@github.com/dbt-labs/internal-package.git"
```

{% hint style="info" %}
**Environment Variables**

When using environment variables with dbt, ensure they're available in your execution environment. You can set these as environment variables in your operating system or in your CI/CD pipeline.
{% endhint %}

**SSH Key Method (Command Line)**

For command-line users with SSH authentication:

```yaml
packages:
  - git: "git@github.com:dbt-labs/internal-package.git"
```

***

### Package Maintenance

**Updating Packages**

To update packages:

1. Change the version/revision in `packages.yml`
2. Run `dbt deps` to install the updated packages
3. Test the changes thoroughly before deploying

**Uninstalling Packages**

To remove a package:

1. Delete it from your `packages.yml` file
2. Run `dbt clean` to remove the installed package
3. Run `dbt deps` to reinstall remaining packages

***

#### Advanced Package Techniques

**Handling Package Conflicts**

When using multiple packages, you might encounter naming conflicts. You can resolve these by:

1. Using fully-qualified references:

   ```sql
   {{ dbt_utils.generate_surrogate_key(['id']) }}
   ```
2. Overriding package macros in your project:

   ```sql
   {% macro generate_surrogate_key(field_list) %}
       {# Your custom implementation #}
   {% endmacro %}
   ```

**Subdirectory Configuration**

For packages nested in subdirectories (e.g., in monorepos):

```yaml
packages:
  - git: "https://github.com/dbt-labs/dbt-labs-experimental-features"
    subdirectory: "materialized-views"
```


# Model Materializations

Covers dbt's model materialization strategies, including table, view, incremental, and ephemeral, and how to configure materializations at the project and individual model level.

Materializations are strategies that determine how dbt™ persists your models in the data warehouse. Think of them as different ways to store and update your transformed data.

### Available Materialization Types

| Type                                                                                                  | Description                                | Best For                                     |
| ----------------------------------------------------------------------------------------------------- | ------------------------------------------ | -------------------------------------------- |
| [View](/app-help/concepts/dbt-fundamentals/model-materializations/view-materialization)               | A saved query that runs on-demand          | Simple transformations, real-time data needs |
| [Table](/app-help/concepts/dbt-fundamentals/model-materializations/table-materialization)             | A physically stored copy of your data      | BI tools, complex queries, frequent access   |
| [Incremental](/app-help/concepts/dbt-fundamentals/model-materializations/incremental-materialization) | A table that updates only new/changed data | Large datasets, frequent updates             |
| [Ephemeral](/app-help/concepts/dbt-fundamentals/model-materializations/ephemeral-materialization)     | Code that's injected into dependent models | Large datasets, frequent updates             |

{% hint style="info" %}
If you don't specify a materialization, dbt™ will create your model as a [view](/app-help/concepts/dbt-fundamentals/model-materializations/view-materialization) by default.
{% endhint %}

***

### Configuring Materializations

You can configure materializations in two ways:

#### 1. Project Level (dbt\_project.yml)

{% code title="dbt\_project.yml" %}

```yaml
models:
  your_project:
    finance:
      +materialized: table    # All finance models as tables
    staging:
      +materialized: view     # All staging models as views
```

{% endcode %}

#### 2. Model Level (in .sql files)

```sql
{{ 
  config(
    materialized='table'
  )
}}

select * from ...
```

***

### Choosing the Right Materialization

Consider these factors when selecting a materialization:

| Factor              | Consideration                              |
| ------------------- | ------------------------------------------ |
| 🔄 Data Freshness   | How current does the data need to be?      |
| ⚡ Query Performance | How current does the data need to be?      |
| 📊 Data Volume      | How much data are you transforming?        |
| 💰 Resource Cost    | What are your compute/storage constraints? |

***

{% hint style="info" %}
Learn more about each materialization type in their dedicated documentation pages:

* [View Materialization](/app-help/concepts/dbt-fundamentals/model-materializations/view-materialization)
* [Table Materialization](/app-help/concepts/dbt-fundamentals/model-materializations/table-materialization)
* [Incremental Materialization](/app-help/concepts/dbt-fundamentals/model-materializations/incremental-materialization)
* [Ephemeral Materialization](/app-help/concepts/dbt-fundamentals/model-materializations/ephemeral-materialization)
  {% endhint %}


# Table Materialization

Learn about table materializations in dbt, understand when to use them, and explore their advantages and limitations.

## Table Materialization

A table materialization rebuilds your model as a physical table in your data warehouse during each dbt run. Unlike views, tables store the actual data rather than just the query logic, using a `CREATE TABLE AS` statement.

***

### How Table Materializations Work

When you materialize a model as a table, dbt executes the model's SQL query and stores the results as a table in your data warehouse. During each run, dbt:

1. Runs `DROP TABLE IF EXISTS` on the existing table
2. Executes `CREATE TABLE AS` with your model's SQL
3. Applies any configured table properties (like indexes, distribution keys, etc.)

This means:

* Data is physically stored in your warehouse
* Queries against the table are faster but builds take longer
* Data isn't automatically updated when source data changes

Under the hood, dbt executes a `CREATE TABLE AS` statement:

```sql
CREATE OR REPLACE TABLE "database"."schema"."customer_orders" AS (
    SELECT * FROM source_table WHERE condition = true
);
```

{% hint style="info" %}
Tables are ideal when query performance is more important than build time or real-time data needs.
{% endhint %}

***

### When to Use Table Materializations

Tables are particularly valuable for:

| Use Case                    | Why Tables Work Well                                                           |
| --------------------------- | ------------------------------------------------------------------------------ |
| Performance-critical models | Tables provide the fastest query performance, ideal for dashboards and reports |
| Complex transformations     | Compute-intensive operations only need to run once during build                |
| Frequently accessed data    | Multiple users or systems can query without recomputing                        |
| Downstream dependencies     | When many models reference this data, tables reduce overall processing         |

***

### Configuring Table Materializations

Tables can be configured at both the model and project level.

#### Model-Level Configuration

```sql
-- In your model SQL file
{{
    config(
        materialized='table',
        sort='order_date',
        dist='customer_id'
    )
}}

SELECT
    customer_id,
    order_date,
    SUM(amount) as total_amount
FROM {{ ref('stg_orders') }}
GROUP BY 1, 2
```

#### Project-Level Configuration

```yaml
# In your dbt_project.yml file
models:
  your_project:
    marts:
      +materialized: table
```

This sets all models in the `marts/` directory to materialize as tables.

***

### Performance Optimization

Different warehouses offer specific optimization options for tables:

| Warehouse | Optimization Options                                                                |
| --------- | ----------------------------------------------------------------------------------- |
| Snowflake | <p>• Clustering keys<br>• Search optimization<br>• Automatic query optimization</p> |
| BigQuery  | <p>• Partitioning<br>• Clustering<br>• Table expiration</p>                         |
| Redshift  | <p>• Distribution keys<br>• Sort keys<br>• Table compression</p>                    |

To apply these optimizations, use the `config()` function with warehouse-specific parameters:

```sql
{{
    config(
        materialized='table',
        snowflake_cluster_by=['customer_id', 'order_date'],
        bigquery_partition_by={
            "field": "order_date",
            "data_type": "date"
        }
    )
}}
```

***

### Advantages and Limitations

| Advantages                   | Limitations               |
| ---------------------------- | ------------------------- |
| ⚡ Fast query performance     | 🕒 Slower build times     |
| 📊 Efficient for BI tools    | 🔄 No automatic updates   |
| 💪 Great for complex queries | 💾 Uses more storage      |
| 🔀 Ideal for multiple users  | 📈 Higher warehouse costs |

***

### When to Consider Other Materializations

While tables are powerful, consider alternatives when:

* Data needs to be real-time (use views)
* Table is very large and only needs incremental updates (use incremental)
* Model is a simple intermediate transformation used by only one downstream model (use ephemeral)

***

### Best Practices

1. **Materialization Strategy**: Use tables for final reporting layers and complex transformations
2. **Build Frequency**: Schedule table rebuilds based on source data update frequency
3. **Performance Tuning**: Apply appropriate indexes, partitioning, or clustering for your warehouse
4. **Resource Management**: Schedule builds during off-peak hours for large tables
5. **Monitoring**: Track build times and storage usage to identify optimization opportunities

By using table materializations strategically, you can balance performance needs with resource utilization to create an efficient data transformation pipeline.


# View​ Materialization

Learn about view materializations in dbt, understand when to use them, and explore their advantages and limitations.

A view materialization creates a view in your data warehouse that represents the SQL query of your dbt model. Unlike tables, views don't store data physically – they're simply stored query definitions that run each time they're accessed.

***

### How View Materializations Work

When you materialize a model as a view, dbt creates or replaces a view in your warehouse. During each run, dbt:

1. Creates or replaces the view definition using your model's SQL
2. Stores the query definition, not the actual data
3. When queried later, the view executes its underlying SQL on-demand

This means:

* No physical data storage – just the query definition
* Data is always up-to-date with source changes
* Queries run the entire transformation each time
* Build times are faster since no data is materialized

Under the hood, dbt executes a `CREATE VIEW` or `CREATE OR REPLACE VIEW` statement:

```sql
CREATE OR REPLACE VIEW "database"."schema"."my_view" AS (
    SELECT * FROM source_table WHERE condition = true
);
```

{% hint style="info" %}
Views are ideal when you need real-time data or when build time is more important than query performance.
{% endhint %}

***

### When to Use View Materializations

Views are particularly valuable for:

| Use Case                   | Why Views Work Well                                  |
| -------------------------- | ---------------------------------------------------- |
| Real-time data needs       | Views always reflect the latest source data          |
| Staging/simple models      | Low-complexity transformations perform well as views |
| Infrequently accessed data | Minimizes storage costs for rarely-used data         |
| Rapid development          | Quick iteration cycle during development             |

***

### Configuring View Materializations

Views can be configured at both the model and project level.

#### Model-Level Configuration

```sql
-- In your model SQL file
{{
    config(
        materialized='view'
    )
}}

SELECT
    customer_id,
    first_name,
    last_name,
    email
FROM {{ ref('stg_customers') }}
WHERE status = 'active'
```

#### Project-Level Configuration

```yaml
# In your dbt_project.yml file
models:
  your_project:
    staging:
      +materialized: view
```

This sets all models in the `staging/` directory to materialize as views.

***

### Performance Considerations

Views have different performance characteristics across warehouses:

| Warehouse | View Characteristics                                                                   |
| --------- | -------------------------------------------------------------------------------------- |
| Snowflake | <p>• Secure views option<br>• Materialized views available<br>• Query optimization</p> |
| BigQuery  | <p>• Authorized views<br>• Materialized views<br>• Query caching</p>                   |
| Redshift  | <p>• Late binding views<br>• Materialized views<br>• Query planning</p>                |

To apply specific view configurations, use the `config()` function with appropriate parameters:

```sql
{{
    config(
        materialized='view',
        secure=true,
        bind=false
    )
}}
```

***

### Advantages and Limitations

| Advantages                        | Limitations                                  |
| --------------------------------- | -------------------------------------------- |
| ⚡ Fast build times                | 🐢 Slower query performance                  |
| 🔄 Always reflects current data   | ⚠️ Resource-intensive for complex queries    |
| 💾 Minimal storage usage          | ⏱️ Each query recomputes the transformation  |
| 🔍 Shows exact lineage in queries | 📊 Can create performance issues in BI tools |

***

### When to Consider Other Materializations

While views are powerful, consider alternatives when:

* Query performance becomes critical (use tables)
* Transformations are complex and compute-intensive (use tables)
* View references lots of data but users only need recent records (use incremental models)
* Transformation is only a stepping stone for a single downstream model (use ephemeral)

***

### Best Practices

1. **Default to Views**: Start with views for most models and change only when needed
2. **Staging Models**: Keep staging models as views for flexibility
3. **Query Optimization**: Write efficient SQL to reduce runtime overhead
4. **Monitor Performance**: Watch for slow-running views and consider materializing as tables
5. **Documentation**: Clearly document performance expectations for view models

By using view materializations strategically, you can create flexible, always-up-to-date data transformations while minimizing storage costs and build times.


# Incremental Materialization

Explains dbt's incremental modeling capabilities for updating only   new/modified data in large/complex transformations, covering different   strategies and best practices.

Incremental models allow you to update only new or modified data in your warehouse instead of rebuilding entire tables. This optimization is particularly valuable when working with:

* Large datasets (millions/billions of rows)
* Computationally expensive transformations
* Time-series data with frequent updates

### Basic Configuration

{% hint style="info" %}
While these examples use Snowflake syntax, the core concepts apply to most data warehouses. Specific syntax and available features may vary by platform.
{% endhint %}

```sql
{{
  config(
    materialized='incremental',
    unique_key='id'
  )
}}

SELECT 
    id,
    status,
    amount,
    updated_at
FROM {{ ref('stg_source') }}

{% if is_incremental() %}
    -- This filter will only be applied on an incremental run
    WHERE updated_at > (SELECT MAX(updated_at) FROM {{ this }})
{% endif %}
```

**Key Components**

1. **Materialization Config**: Set `materialized='incremental'` in your config block
2. **Unique Key**: Define what makes each row unique (single column or multiple columns)
3. **Incremental Logic**: Use the `is_incremental()` macro to filter for new/changed records

***

### Incremental Strategies

dbt™ supports several strategies for incremental models, each with specific use cases:

| Strategy                                                                                                                                            | Description                                            | Best For                                  |
| --------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------ | ----------------------------------------- |
| [Merge](/app-help/concepts/dbt-fundamentals/model-materializations/incremental-materialization/using-merge-for-incremental-models)                  | Updates existing records and inserts new ones          | Tables requiring both inserts and updates |
| [Delete+Insert](/app-help/concepts/dbt-fundamentals/model-materializations/incremental-materialization/using-delete+insert-for-incremental-models)  | Deletes matching records and reinserts new versions    | Batch updates where most records change   |
| [Append-Only](/app-help/concepts/dbt-fundamentals/model-materializations/incremental-materialization/using-append-for-incremental-models)           | Simply adds new records without updating existing ones | Event logs, immutable data                |
| [Microbatch (Beta)](/app-help/concepts/dbt-fundamentals/model-materializations/incremental-materialization/using-microbatch-for-incremental-models) | Processes updates in smaller batches                   | Very large time-series datasets           |

For detailed examples and configuration options for each strategy, see their dedicated pages.

***

### Advanced Features

**1. Schema Change Management**

Handle column additions or removals with the `on_schema_change` parameter:

```sql
{{
  config(
    materialized='incremental',
    on_schema_change='sync_all_columns'  -- Options: ignore, fail, append_new_columns, sync_all_columns
  )
}}
```

Options explained:

* `sync_all_columns`: Automatically adapts to column changes (recommended)
* `fail`: Halts execution when schema changes (useful during development)
* `ignore`: Maintains existing schema (use cautiously)
* `append_new_columns`: Adds new columns without removing old ones

**2. Incremental Predicates**

Optimize performance for large datasets:

```sql
{{
  config(
    materialized='incremental',
    incremental_predicates=[
      "DBT_INTERNAL_DEST.session_start > dateadd(day, -7, current_date)"
    ],
    cluster_by=['session_start']
  )
}}
```

This configuration:

* Limits the scan of existing data
* Improves merge performance
* Works with clustering for better query optimization

**3. Strategy-Specific Configurations**

Control column updates in merge operations:

```sql
{{
  config(
    materialized='incremental',
    merge_update_columns=['email', 'ip_address'],  -- Only update these columns
    merge_exclude_columns=['created_at']  -- Never update these columns
  )
}}
```

**4. Custom Strategies**

Create your own incremental strategy:

```sql
-- macros/my_custom_strategies.sql
{% macro get_incremental_insert_only_sql(arg_dict) %}
  {% do return(some_custom_macro_with_sql(
    arg_dict["target_relation"],
    arg_dict["temp_relation"],
    arg_dict["unique_key"],
    arg_dict["dest_columns"],
    arg_dict["incremental_predicates"]
  )) %}
{% endmacro %}

-- models/my_model.sql
{{
  config(
    materialized='incremental',
    incremental_strategy='insert_only'
  )
}}
```

***

### Best Practices

**1. Handle Late-Arriving Data**

Data doesn't always arrive in perfect chronological order. Include a buffer period in your incremental logic:

```sql
{% if is_incremental() %}
    -- Look back 3 days to catch late-arriving data
    WHERE updated_at > DATEADD(days, -3, (SELECT MAX(updated_at) FROM {{ this }}))
{% endif %}
```

**2. Optimize Performance**

Use appropriate configurations to improve query performance and efficiency:

```sql
{{
  config(
    materialized='incremental',
    cluster_by=['date_field', 'customer_id'],  -- Improve query performance
    merge_update_columns=['status', 'amount']  -- Minimize update overhead
  )
}}
```

**3. Regular Maintenance**

To prevent potential data inconsistencies that might accumulate over time, periodically rebuild your entire table:

```bash
dbt run --full-refresh --select model_name
```

**4. Multiple Column Keys**

When a single column isn't enough to identify unique records:

```sql
{{
  config(
    materialized='incremental',
    unique_key=['customer_id', 'order_date']
  )
}}
```


# Using Merge for Incremental Models

Details the merge strategy for incremental models, which updates existing records   and inserts new ones using a database MERGE statement.

The merge method uses your database's MERGE statement (or equivalent) to update existing records and insert new ones. This method provides the most control over how your incremental model is updated but requires more computational resources.

***

### When to Use Merge Method

* Tables requiring both inserts and updates
* When data consistency is critical
* When specific columns need updating
* For maintaining referential integrity

***

### Advantages and Trade-offs

| Advantages                          | Trade-offs                            |
| ----------------------------------- | ------------------------------------- |
| Most flexible method                | Slower than other methods             |
| Precise control over column updates | More resource-intensive               |
| Maintains data consistency          | Can be costly for very large datasets |
| Handles complex update patterns     | Requires a unique key                 |

***

### Example Implementation

```sql
{{
  config(
    materialized='incremental',
    unique_key='event_id',
    incremental_strategy='merge'  // Default, can be omitted
  )
}}

SELECT 
    event_id,
    event_type,
    event_time,
    user_id,
    properties
FROM {{ ref('stg_events') }}

{% if is_incremental() %}
    WHERE event_time > (SELECT MAX(event_time) FROM {{ this }})
{% endif %}
```

In this example:

* `unique_key` identifies which records should be updated
* The `WHERE` clause in the `is_incremental()` block filters for new records
* Existing records with matching `event_id` values will be updated
* New records will be inserted

***

### Fine-Tuning Merge Operations

You can precisely control which columns are updated during merge operations:

```sql
{{
  config(
    materialized='incremental',
    unique_key='user_id',
    incremental_strategy='merge',
    merge_update_columns=['email', 'last_login_at', 'profile_updated_at'],  -- Only update these columns
    merge_exclude_columns=['created_at', 'signup_source']  -- Never update these columns
  )
}}
```

This configuration gives you fine-grained control over the merge process.

***

### Step By Step Guide: Merge Method Implementation

Let's see how the merge method works with a practical example.

**Step 1: Initial Setup and Testing**

First, create the staging model:

```sql
-- models/staging/stg_test_source.sql
WITH source_data AS (
    SELECT 1 as id,
           'active' as status,
           100 as amount,
           CURRENT_TIMESTAMP() as updated_at
    UNION ALL
    SELECT 2, 'pending', 200, CURRENT_TIMESTAMP()
    UNION ALL
    SELECT 3, 'active', 300, CURRENT_TIMESTAMP()
)
SELECT * FROM source_data
```

Run staging model:

```bash
dbt run --select stg_test_source
```

Verify staging data in Snowflake:

```sql
SELECT 
    * 
FROM 
    your_database.your_schema.stg_test_source 
ORDER BY 
    id
```

Expected Result:

```
ID | STATUS  | AMOUNT | UPDATED_AT
1  | active  | 100    | [timestamp]
2  | pending | 200    | [timestamp]
3  | active  | 300    | [timestamp]
```

Create the incremental model:

```sql
-- models/incremental_test.sql
{{
  config(
    materialized='incremental',
    unique_key='id'
  )
}}

SELECT 
    id,
    status,
    amount,
    updated_at
FROM {{ ref('stg_test_source') }}

{% raw %}
{% if is_incremental() %}
    WHERE updated_at > (SELECT MAX(updated_at) FROM {{ this }})
{% endif %}
{% endraw %}
```

Run incremental model:

```bash
dbt run --select incremental_test
```

The result should match the staging model (3 rows).

**Step 2: Testing Merge Method**

Update staging model with changes:

```sql
-- models/staging/stg_test_source.sql
WITH source_data AS (
    SELECT 1 as id,
           'inactive' as status,  -- Changed status
           150 as amount,         -- Changed amount
           CURRENT_TIMESTAMP() as updated_at
    UNION ALL
    SELECT 4, 'new', 400, CURRENT_TIMESTAMP()  -- New record
)
SELECT * FROM source_data
```

Run staging model:

```bash
dbt run --select stg_test_source
```

Verify staging data:

```sql
SELECT 
    * 
FROM 
    your_database.your_schema.stg_test_source 
ORDER BY 
    id;
```

Expected Result:

```
ID | STATUS   | AMOUNT | UPDATED_AT
1  | inactive | 150    | [new_timestamp]
4  | new      | 400    | [new_timestamp]
```

Run incremental model:

```bash
dbt run --select incremental_test
```

Verify incremental model:

```sql
SELECT 
    * 
FROM 
    your_database.your_schema.incremental_test 
ORDER BY 
    id
```

Expected Result:

```
ID | STATUS   | AMOUNT | UPDATED_AT
1  | inactive | 150    | [new_timestamp]  -- Updated
2  | pending  | 200    | [old_timestamp]  -- Unchanged
3  | active   | 300    | [old_timestamp]  -- Unchanged
4  | new      | 400    | [new_timestamp]  -- New
```

Notice that record with ID 1 was updated with the new status and amount, while new record with ID 4 was inserted. Records with IDs 2 and 3 remained unchanged as they weren't in the staging data.

***

### Common Issues and Solutions

<table><thead><tr><th width="256.17578125">Issue</th><th>Solution</th></tr></thead><tbody><tr><td>Merge performance with large tables</td><td>Use <code>incremental_predicates</code> to limit the scope of the merge operation</td></tr><tr><td>Unexpected column updates</td><td>Use <code>merge_update_columns</code> or <code>merge_exclude_columns</code> for precise control</td></tr><tr><td>Duplicate key errors</td><td>Ensure your <code>unique_key</code> truly identifies records uniquely</td></tr></tbody></table>

The merge method offers the most flexibility for incremental models, but comes with higher computational costs. Consider it your default choice unless you have specific performance requirements or data characteristics that favor another method.


# Using Delete+Insert for Incremental Models

Covers the delete+insert method for incremental models, which replaces batches    of records by first deleting them and then inserting new versions.

The delete+insert method is like a "replace-all" approach for updated records. It first removes all records that match your incremental criteria and then inserts the new versions of those records. Think of it as replacing an entire page in a book rather than editing individual words.

***

### When to Use Delete+Insert Method

* Batch updates where most records change
* When merge performance becomes a bottleneck
* Simpler change patterns
* Large-scale updates
* When most fields in each record need updating

***

### Advantages and Trade-offs

| Advantages                                     | Trade-offs                                              |
| ---------------------------------------------- | ------------------------------------------------------- |
| Better performance than merge for bulk updates | Less granular control than merge                        |
| Simpler execution plan                         | All columns are updated (no column-specific updates)    |
| Good for high-volume changes                   | Potential for higher resource usage during delete phase |
| Simpler SQL generated                          | Doesn't work well for individual record updates         |

#### Example Implementation

```sql
{{
  config(
    materialized='incremental',
    incremental_strategy='delete+insert',
    unique_key='id'
  )
}}

SELECT 
    id,
    status,
    batch_date,
    amount,
    updated_at
FROM {{ ref('stg_source') }}

{% if is_incremental() %}
    WHERE batch_date >= (SELECT MAX(batch_date) FROM {{ this }})
{% endif %}
```

In this example:

* Records matching the `WHERE` condition will first be deleted
* Then all new/updated records will be inserted
* This is more efficient than merge when most fields need updating

***

### How It Works

When you run a delete+insert incremental model, dbt:

1. Creates a temporary table with all the data from your query
2. Identifies records in the target table that match the incremental filter condition
3. Deletes those matching records from the target table
4. Inserts all records from the temporary table into the target table

The result is similar to a merge, but with a different execution strategy that can be more efficient in specific scenarios.

***

### Step By Step Guide: Delete+Insert Method Implementation

Let's see how the delete+insert method works with a practical example.

**Initial Setup**

Follow the same initial setup as the [Merge Method example](/app-help/concepts/dbt-fundamentals/model-materializations/incremental-materialization/using-merge-for-incremental-models#step-by-step-guide-merge-method-implementation), creating the staging model and first version of the incremental model.

**Testing Delete+Insert Method**

1. Update incremental model config (keep same staging data):

```sql
-- models/incremental_test.sql
{{
  config(
    materialized='incremental',
    incremental_strategy='delete+insert',
    unique_key='id'
  )
}}

SELECT 
    id,
    status,
    amount,
    updated_at
FROM {{ ref('stg_test_source') }}

{% if is_incremental() %}
    WHERE updated_at > (SELECT MAX(updated_at) FROM {{ this }})
{% endif %}
```

2. Update staging model with changes:

```sql
-- models/staging/stg_test_source.sql
WITH source_data AS (
    SELECT 1 as id,
           'inactive' as status,  -- Changed status
           150 as amount,         -- Changed amount
           CURRENT_TIMESTAMP() as updated_at
    UNION ALL
    SELECT 2 as id,
           'approved' as status,  -- Changed from 'pending'
           200 as amount,
           CURRENT_TIMESTAMP() as updated_at
    UNION ALL
    SELECT 4, 'new', 400, CURRENT_TIMESTAMP()  -- New record
)
SELECT * FROM source_data
```

3. Run staging model:

```bash
dbt run --select stg_test_source
```

4. Run incremental model:

```bash
dbt run --select incremental_test
```

5. Verify in Snowflake:

```sql
SELECT 
    * 
FROM 
    your_database.your_schema.incremental_test 
ORDER BY 
    id
```

Expected Result:

```
ID | STATUS   | AMOUNT | UPDATED_AT
1  | inactive | 150    | [new_timestamp]  -- Updated
2  | approved | 200    | [new_timestamp]  -- Updated with new status
3  | active   | 300    | [old_timestamp]  -- Unchanged (wasn't in staging)
4  | new      | 400    | [new_timestamp]  -- New
```

Notice that all records in the staging model (IDs 1, 2, and 4) were processed as a batch - deleted and reinserted. Record 3 remained unchanged as it wasn't in the staging data and didn't match the incremental condition.

***

### Best Practices for Delete+Insert

1. **Use with Batch Processing** - This method works best when processing logical batches of data (like daily updates)
2. **Optimize the Filter Condition** - Choose an incremental filter that selects all records that might need updating but minimizes unnecessary processing
3. **Consider Performance Impact** - The delete operation affects indexes and can cause fragmentation; plan your maintenance accordingly
4. **Pair with Partitioning** - If your database supports it, partitioning can make delete+insert operations more efficient

The delete+insert method is particularly effective for batch-oriented workflows where most record fields change together, striking a balance between the flexibility of merge and the performance of append-only approaches.


# Using Append for Incremental Models

Explains the append-only method for incremental models, ideal for event logs   and immutable data where records are only added, never updated.

The append method is the simplest incremental approach - it just adds new records to your table without updating existing ones. Think of it as adding new entries to a log book where historical entries are never changed. This method is perfect for event logging and similar use cases where historical data remains unchanged.

***

### When to Use Append Method

* Event logs
* Immutable data
* Time-series data without updates
* When duplicates are acceptable
* Audit trails
* High-volume data capture with minimal processing

***

### Advantages and Trade-offs

| Advantages              | Trade-offs                                    |
| ----------------------- | --------------------------------------------- |
| Fastest performance     | No update capability                          |
| Simplest execution plan | Can create duplicate data                     |
| Minimal resource usage  | Requires downstream deduplication if needed   |
| No need for unique key  | Not suitable for dimensions or reference data |
| Scales extremely well   | Can lead to table size management challenges  |

### Example Implementation

```sql
{{
  config(
    materialized='incremental',
    incremental_strategy='append'
  )
}}

SELECT 
    event_id,
    event_type,
    event_timestamp,
    user_id,
    page_url,
    properties
FROM {{ ref('stg_events') }}

{% if is_incremental() %}
    WHERE event_timestamp > (SELECT MAX(event_timestamp) FROM {{ this }})
{% endif %}
```

In this example:

* New records are simply appended to the table
* No updates are made to existing records
* Note that `unique_key` is not required

***

### How It Works

When you run an append-only incremental model, dbt:

1. Creates a temporary table with all the data from your query
2. Inserts all records from the temporary table into the target table
3. No deletion or updating of existing records occurs

This is the simplest and fastest approach, making it ideal for high-volume event data.

***

### Step By Step Guide: Delete+Insert Method Implementation

Let's see how the delete+insert method works with a practical example.

**Initial Setup**

First, follow the same [initial setup](/app-help/concepts/dbt-fundamentals/model-materializations/incremental-materialization/using-merge-for-incremental-models#step-by-step-guide-merge-method-implementation) as described in the "Using Merge for Incremental Models" page:

1. Create the staging model with test data
2. Run the staging model
3. Create the initial incremental model
4. Run the incremental model

Once you have this setup in place, proceed with testing the delete+insert method:

#### Testing Append Method

1. Update incremental model config:

```sql
-- models/incremental_test.sql
{{
  config(
    materialized='incremental',
    incremental_strategy='append'
  )
}}

SELECT 
    id,
    status,
    amount,
    updated_at
FROM {{ ref('stg_test_source') }}

{% if is_incremental() %}
    WHERE updated_at > (SELECT MAX(updated_at) FROM {{ this }})
{% endif %}
```

2. Run incremental model with full refresh to start clean:

```bash
dbt run --full-refresh --select incremental_test
```

3. Verify initial state:

```sql
SELECT 
    * 
FROM 
    your_database.your_schema.incremental_test 
ORDER BY 
    id
```

4. Update staging data to trigger append:

```sql
-- models/staging/stg_test_source.sql
WITH source_data AS (
    SELECT 1 as id,
           'inactive' as status,
           150 as amount,
           CURRENT_TIMESTAMP() as updated_at
    UNION ALL
    SELECT 5, 'newest', 500, CURRENT_TIMESTAMP()  -- Another new record
)
SELECT * FROM source_data
```

5. Run both models:

```bash
dbt run --select stg_test_source
dbt run --select incremental_test
```

6. Verify append results:

```sql
SELECT 
    * 
FROM 
    your_database.your_schema.incremental_test 
ORDER BY 
    updated_at
```

Expected Result:

```
ID | STATUS   | AMOUNT | UPDATED_AT
1  | active   | 100    | [old_timestamp]  -- Original record
2  | pending  | 200    | [old_timestamp]  -- Original record
3  | active   | 300    | [old_timestamp]  -- Original record
1  | inactive | 150    | [new_timestamp]  -- New record (duplicate ID)
5  | newest   | 500    | [new_timestamp]  -- New record
```

Notice that a new record with ID 1 was added, even though it already exists in the table. This is the key characteristic of append-only: it never updates existing records, only adds new ones.

***

### Handling Duplicates with Append Method

Since append can create duplicate records, you may need strategies to handle this downstream:

1. **Window Functions**: Use window functions to identify the most recent version of each record:

```sql
WITH ranked_data AS (
  SELECT 
    *,
    ROW_NUMBER() OVER (PARTITION BY id ORDER BY updated_at DESC) as row_num
  FROM {{ ref('append_only_model') }}
)
SELECT * FROM ranked_data WHERE row_num = 1
```

2. **Materialized Views**: Some warehouses support materialized views that can automatically deduplicate
3. **Downstream Models**: Create downstream models that specifically handle deduplication logic

***

### Best Practices for Append Method

1. **Monitor Table Growth** - Since records are only added, implement a strategy to manage table size
2. **Consider Partitioning** - For very large tables, partitioning by date can improve query performance
3. **Plan for Deduplication** - If uniqueness matters for downstream consumers, build deduplication into your models
4. **Use with Time-Series Data** - Append-only is ideal for time-series data where historical accuracy is important

The append method offers the best performance for high-volume data ingestion where updates to existing records aren't needed, making it the method of choice for event data, logs, and other immutable datasets.


# Using Microbatch for Incremental Models

Explores the microbatch method for incremental models, designed for    processing very large datasets in smaller batches for improved reliability   and performance.

The microbatch method processes incremental updates in smaller batches, designed specifically for very large time-series datasets. It's particularly valuable when dealing with datasets that are too large to process in a single operation.

{% hint style="warning" %}
The microbatch method is currently in Beta. Features and syntax may change in future releases.
{% endhint %}

***

### When to Use Microbatch Method

* Very large time-series datasets
* When reliability is crucial
* When processing needs to be more resilient
* When database resource limits are a concern
* For tables with billions of rows
* When single-transaction operations time out

***

### Advantages and Trade-offs

| Advantages                        | Trade-offs                                          |
| --------------------------------- | --------------------------------------------------- |
| More efficient for large datasets | More complex setup                                  |
| Better error handling             | Requires careful batch size tuning                  |
| Improved resilience               | Only available on certain platforms                 |
| Reduced memory usage              | Slightly more overhead than single-batch operations |
| Can work around query timeouts    | Still in beta status                                |

### Example Implementation

```sql
{{
  config(
    materialized='incremental',
    incremental_strategy='microbatch',
    unique_key='id',
    incremental_predicates=[
      "DBT_INTERNAL_DEST.event_time > dateadd(day, -7, current_date)"
    ]
  )
}}

SELECT 
    id,
    event_type,
    event_time,
    user_id,
    properties
FROM {{ ref('stg_events') }}

{% if is_incremental() %}
    WHERE event_time > (SELECT MAX(event_time) FROM {{ this }})
{% endif %}
```

In this example:

* Updates are processed in smaller batches
* `incremental_predicates` limits the scope of each operation
* The overall process is more resilient to timeouts and failures

***

### How It Works

When you run an incremental model with the microbatch method, dbt:

1. Breaks the update into smaller chunks (batches)
2. Processes each batch in a separate transaction
3. Continues with remaining batches even if some fail
4. Reports overall success or partial failure

This approach is similar to the merge method, but with added reliability for very large datasets.

***

### Configuration Options

The microbatch method has several unique configuration options:

```sql
{{
  config(
    materialized='incremental',
    incremental_strategy='microbatch',
    unique_key='id',
    
    -- Microbatch-specific options
    microbatch_size=5000,  -- Number of rows per batch
    microbatch_limit=250,  -- Maximum number of batches
    incremental_predicates=[
      "DBT_INTERNAL_DEST.created_at > dateadd(day, -7, current_date)"
    ]
  )
}}
```

| Option                   | Description                                 | Default |
| ------------------------ | ------------------------------------------- | ------- |
| `microbatch_size`        | Number of rows per batch                    | 10000   |
| `microbatch_limit`       | Maximum number of batches to process        | 500     |
| `incremental_predicates` | Conditions to limit the scope of each batch | None    |

***

### Best Practices for Microbatch Method

1. **Tune Batch Size** - Find the optimal batch size that balances performance and reliability
2. **Use Incremental Predicates** - Always use predicates to limit the scope of the operation
3. **Monitor Partial Failures** - Set up alerting for when some batches fail but others succeed
4. **Consider Recovery Strategies** - Have plans for reprocessing failed batches

The microbatch method is still evolving but offers a promising solution for handling very large incremental models where other methods might time out or consume too many resources. It's particularly valuable for high-volume event data or IoT datasets where individual transactions might exceed database limitations.


# Ephemeral Materialization

Describes dbt's ephemeral materialization, which integrates models as CTEs in   dependent models, discussing advantages, disadvantages, and recommended use   cases.

## Ephemeral Materialization

Unlike other materializations, ephemeral models don't create database objects. Instead, dbt incorporates the model's code into dependent models using Common Table Expressions (CTEs).

This means:

* No physical tables or views are created
* Code is integrated into downstream models
* Logic can be reused across models

{% hint style="info" %}
Ephemeral models are ideal for lightweight, reusable transformations that don't need to be queried directly.
{% endhint %}

***

#### Working with Ephemeral Models

Let's understand ephemeral models through a simple example:

```sql
-- models/intermediate/clean_customer_data.sql
{{
    config(
        materialized='ephemeral'
    )
}}

SELECT
    customer_id,
    TRIM(first_name) as first_name,
    TRIM(last_name) as last_name,
    LOWER(email) as email
FROM {{ ref('raw_customers') }}
```

When this model is referenced:

1. No table or view is created
2. The code becomes a CTE in downstream models
3. The logic is recomputed each time it's used

***

#### When to Use Ephemeral Models

Ephemeral models work best for:

* **Simple transformations** (cleaning, filtering)
* **Early-stage models** in your DAG
* **Light data preparation** steps
* **Code that supports 1-2 downstream models**

***

#### Configuration

**Model-Level Configuration**

```sql
-- In your model SQL file
{{
    config(
        materialized='ephemeral'
    )
}}

SELECT 
    user_id,
    CASE 
        WHEN age < 18 THEN 'youth'
        WHEN age < 65 THEN 'adult'
        ELSE 'senior'
    END as age_category
FROM {{ ref('raw_users') }}
```

***

#### Performance and Usage Guidelines

**Advantages and Limitations**

<table><thead><tr><th width="338">Advantages</th><th>Limitations</th></tr></thead><tbody><tr><td>🧹 Reduces database clutter</td><td>🚫 Can't query directly</td></tr><tr><td>♻️ Enables code reuse</td><td>⚠️ No model contracts support</td></tr><tr><td>📦 No storage overhead</td><td>🔍 Harder to debug</td></tr><tr><td>🔄 Flexible logic updates</td><td>🛠️ Limited operational use</td></tr></tbody></table>

**When to Change Materialization**

Consider a different materialization when you:

* Need to query the model directly
* Have complex transformations
* Support many downstream models
* Require model contracts

**Best Practices**

1. **Keep It Simple**: Use for lightweight transformations only
2. **Limit Dependencies**: Best for 1-2 downstream models
3. **Consider Debugging**: Use sparingly to maintain query clarity

{% hint style="info" %}
Ephemeral models are powerful for code organization but should be used judiciously. Consider views or tables for more complex or widely-used transformations.
{% endhint %}


# Snapshots

Covers dbt's snapshot functionality for tracking historical changes using Type 2 Slowly Changing Dimensions, including configuration, strategies, metadata, and best practices.

Snapshots track historical changes in your data warehouse by implementing Type 2 Slowly Changing Dimensions (SCD). Instead of overwriting data, snapshots create new records while preserving the history of changes.

#### Understanding Slowly Changing Dimensions

In data warehousing, different strategies exist for handling data that changes over time:

* **Type 0 (Retain Original)**: The dimension never changes once created
* **Type 1 (Overwrite)**: Updates replace the original values with no history kept
* **Type 2 (Add New Row)**: Preserves history by creating new records when changes occur
* **Type 3 (Add New Attribute)**: Maintains limited history by adding columns for previous values

dbt snapshots implement **Type 2 SCD**, which is ideal for tracking changes to data that updates infrequently and requires historical tracking, such as:

* Customer status changes
* Product pricing and categorization
* Order processing states
* Employee roles and departments

Here's an example of how snapshots preserve history when a customer's status changes:

```sql
-- Initial record
customer_id  status    updated_at     dbt_valid_from    dbt_valid_to
101         active     2023-06-15     2023-06-15        null

-- After status change with snapshots
customer_id  status    updated_at     dbt_valid_from    dbt_valid_to
101         active     2023-06-15     2023-06-15        2023-08-20    -- Historical record
101         inactive   2023-08-20     2023-08-20        null          -- Current record
```

Using this snapshot, you can query the customer's status on any specific date.

***

### Basic Configuration

{% hint style="info" %}
**Prerequisites**

Before setting up snapshots:

* Create a target schema for snapshots in your data warehouse
* Ensure source data has either reliable timestamps or columns to track for changes
* Set up the `snapshots/` directory in your dbt project
  {% endhint %}

Snapshots are defined in `.sql` files within your project's `snapshots` directory:

```sql
{% snapshot orders_snapshot %}

{{
    config(
      target_schema='snapshots',
      unique_key='order_id',
      strategy='timestamp',
      updated_at='updated_at'
    )
}}

select * from {{ source('orders', 'orders') }}

{% endsnapshot %}
```

To execute the snapshot, run:

```bash
dbt snapshot
```

This command should be run whenever you want to capture changes, both in development and production environments.

***

### Snapshot Strategies

dbt offers two primary strategies to detect changes in your data:

{% tabs %}
{% tab title="Timestamp Strategy" %}
Uses a timestamp column to detect changes. This is the best choice when your source data reliably updates a timestamp column when records change.

**When to use:**

* Source data has a reliable updated\_at column
* Need to track when changes occurred
* Want to capture all changes based on timing

**Configuration**

```sql
{% snapshot orders_snapshot %}

{{
    config(
      target_schema='snapshots',
      unique_key='order_id',
      strategy='timestamp',
      updated_at='updated_at'
    )
}}

-- Your SQL query

{% endsnapshot %}
```

The updated\_at parameter specifies which column contains the last-updated timestamp.
{% endtab %}

{% tab title="Check Strategy" %}
Detects changes by comparing specific column values. Use when you don't have reliable timestamps or want to track only certain columns.

**When to use:**

* No reliable updated\_at column
* Only want to track specific column changes
* Need control over what constitutes a change

**Configuration**

```sql
{{
    config(
      target_schema='snapshots',
      unique_key='order_id',
      strategy='check',
      check_cols=['status', 'amount']  -- or 'all'
    )
}}
```

The `check_cols` parameter can be:

* A list of specific columns to check for changes
* The string `'all'` to check all columns {% endtab %} {% endtabs %}
  {% endtab %}
  {% endtabs %}

***

### How Snapshots Work Behind the Scenes

When you run `dbt snapshot`, dbt:

1. Creates the snapshot table if it doesn't exist
2. For existing tables, determines which records have changed based on your strategy
3. Marks previously current records that changed as no longer current (`dbt_valid_to` gets timestamp)
4. Inserts new versions of changed records as current (with `dbt_valid_to` as null)
5. Inserts completely new records

This process maintains a complete history of all changes while ensuring current data is easily identifiable.

***

### Metadata Fields

As of dbt version ≥ 1.9, snapshots add four tracking columns:

* `dbt_valid_from`: When this version became valid
* `dbt_valid_to`: When this version became invalid (null for current version)
* `dbt_updated_at`: Timestamp when the snapshot was taken
* `dbt_scd_id`: Unique identifier for each version

These columns allow you to:

* Identify current records (`dbt_valid_to IS NULL`)
* Find records valid at a specific point in time
* Determine when changes occurred
* Track the duration a particular version was active

***

### Step-by-Step Example

Let's implement a simple snapshot to track customer status changes:

**1. Create a staging model for the source data**

```sql
-- models/staging/stg_customers.sql
SELECT
    customer_id,
    name,
    status,
    email,
    updated_at
FROM {{ source('crm', 'customers') }}
```

**2. Create a snapshot file**

```sql
-- snapshots/customer_snapshots.sql
{% snapshot customers_snapshot %}

{{
    config(
      target_schema='snapshots',
      unique_key='customer_id',
      strategy='timestamp',
      updated_at='updated_at'
    )
}}

select * from {{ ref('stg_customers') }}

{% endsnapshot %}
```

**3. Run the snapshot for the first time**

```bash
dbt snapshot
```

This creates the initial snapshot table with all customers.

**4. Simulate a data change in the source**

When customer data changes in your source system, run the snapshot again to capture those changes:

```bash
dbt snapshot
```

**5. Query historical and current data**

```sql
-- Get current customer data
SELECT * FROM snapshots.customers_snapshot
WHERE dbt_valid_to IS NULL;

-- Get customer data as it existed on a specific date
SELECT * FROM snapshots.customers_snapshot
WHERE customer_id = 123
  AND dbt_valid_from <= '2023-06-01'
  AND (dbt_valid_to > '2023-06-01' OR dbt_valid_to IS NULL);
```

***

### Advanced Configuration

**Invalidating Hard Deletes**

By default, snapshots don't track when records are deleted from the source. To track deletions:

```sql
{% snapshot orders_snapshot %}

{{
    config(
        target_schema='snapshots',
        unique_key='order_id',
        strategy='timestamp',
        updated_at='updated_at',
        invalidate_hard_deletes=true
    )
}}

-- Your SQL query

{% endsnapshot %}
```

With `invalidate_hard_deletes=true`, dbt will:

* Identify records in the snapshot that no longer exist in the source
* Set their `dbt_valid_to` timestamps to mark them as no longer current

**Custom Snapshot Schemas**

You can dynamically set the target schema:

```sql
{% snapshot orders_snapshot %}

{{
    config(
        target_schema=var('snapshot_schema', 'snapshots'),
        unique_key='order_id',
        strategy='timestamp',
        updated_at='updated_at'
    )
}}

-- Your SQL query

{% endsnapshot %}
```

This allows you to use variables to change the schema at runtime.

***

### Best Practices for Snapshots

**1. Snapshot Frequency**

* Schedule snapshots based on how frequently your source data changes and how important it is to capture every change
* For critical data, run snapshots before dependent models to ensure they use the latest history
* Consider storage costs versus historical data needs

```yaml
# Example schedule.yml for a daily snapshot
jobs:
  - name: daily_snapshot
    schedule: "0 1 * * *"  # 1:00 AM daily
    steps:
      - dbt snapshot
      - dbt run --models dependent_on_snapshots
```

**2. Performance Optimization**

* Use a dedicated schema for snapshots to make maintenance easier
* Apply appropriate indexes to snapshot tables for faster querying
* Consider partitioning large snapshot tables by date
* For very large tables, use incremental snapshots with time-based filters

**3. Querying Historical Data**

To access historical data, filter to see the state of a specific record as it existed on a given date:

```sql
-- Get data as it looked on a specific date
SELECT *
FROM {{ ref('customers_snapshot') }}
WHERE customer_id = 123
  AND dbt_valid_from <= '2023-06-01'
  AND (dbt_valid_to > '2023-06-01' OR dbt_valid_to IS NULL)
```

**4. Common Pitfalls to Avoid**

* **Infrequent snapshots**: Running snapshots too infrequently might miss intermediate state changes
* **Missing source filter**: Always filter your source query to include only necessary data
* **Unreliable timestamps**: Ensure your `updated_at` field actually updates when records change
* **Using snapshots for high-frequency changes**: Consider incremental models for data that changes very frequently

***

### Using Snapshots in Your Data Architecture

Snapshots typically fit into your data architecture as follows:

```
dbt project
├── models/
│   ├── staging/         # Simple transformations of source data
│   ├── intermediate/    # Business logic transformations
│   └── marts/           # Business-level output models
├── snapshots/           # Historical tracking of changing data
│   ├── customer_snapshots.sql
│   └── product_snapshots.sql
└── analyses/            # SQL for historical analysis using snapshots
    └── customer_status_history.sql
```

By effectively implementing snapshots, you create a robust history tracking system that supports historical reporting, audit requirements, and trend analysis—all while maintaining the simplicity and reproducibility that makes dbt powerful.


# Running dbt™


# Mastering the dbt™ CLI


# Commands

Covers the key dbt CLI commands and their use cases, including run, test, source freshness, compile, documentation, and utility commands to control data transformations.

The dbt™ CLI offers a range of commands for executing data transformations. Each command has its own options and parameters, allowing you to precisely control your data transformations. Let's explore these commands and their common use cases.

<figure><img src="/files/P5k9nCkVKnCZpNuQa6Ha" alt=""><figcaption></figcaption></figure>

### The Basics: dbt run

The bread and butter of dbt™ is the run command. It's like hitting the "Go" button on your data transformations. The dbt run command is the most complex and can be broken down into 4 parts:

* **Arguments** like --select, --exclude and others
* **Model names** to choose what models to run
* **Method selectors** offering ability to fine tune which models to run
* **Graph selectors** offering further fine tuning to apply complex boolean-like logic

<figure><img src="/files/f6tEc7NHdheyYLCBAoGR" alt=""><figcaption></figcaption></figure>

Further configure your `dbt run` command with these options:

```bash
# Run specific models
dbt run --select cool_waffle

# Skip certain models
dbt run --exclude boring_jaffle

# Rebuild everything from scratch
dbt run --full-refresh

# Pass variables to models
dbt run --vars '{"my_var": "value"}'

# Speed up runs with multiple threads
dbt run --threads 4
```

### Running Tests

Don't let bad data crash your party.

Use dbt test to keep your transformations in check and apply data quality best practices to your dbt™ transformation pipelines:

```bash
dbt test

# Test specific models
dbt test --select critical_data

# Run schema tests only
dbt test --select "test_type:generic"
```

### Source Freshness

Source freshness in dbt™ is like a built-in data freshness checker. It helps you:

* Monitor when your source data was last updated
* Set expectations for how recent your data should be
* Alert you when data is stale

To check the freshness of all your defined sources, run:

```bash
dbt source freshness
```

### Compile

Use dbt compile to convert all your dbt™ models with their Jinja references into raw SQL. This is the SQL dbt™ will run against your data warehouse. It's like X-ray vision for your SQL:

```bash
bashCopydbt compile
```

When your dbt™ models fail to run, you need to start with the compiled SQL first.

### Generate Documentation

Convert all your schema and table descriptions into static HTML files and then serve them from a server or cloud bucket like AWS S3.

```bash
dbt docs generate
dbt docs serve
```

### Debug Mode

When you can't make head or tail of errors you're seeing during development or production runs, use the --debug option. This will generate additional logs in your terminal to help triage the situation. This is most useful in diagnosing warehouse connection errors.

```bash
dbt run --debug
```

### The Snapshot

Capture data changes over time:

```bash
dbt snapshot
```

### Build Everything

The all-in-one command for the impatient:

```bash
dbt build
```

It runs, tests, and snapshots in one go.

### CSVs: dbt seed

Convert CSV files to tables:

```bash
dbt seed
```

### List Models: dbt ls

List your models:

```bash
dbt ls

# List the most important resources
dbt ls --select tag:important
```

### Preview Model Output: dbt show

Preview your model's output:

```bash
dbt show --select cool_waffle
```

### Retry When Something Fails

Oops, something failed? Try again:

```bash
dbt retry
```

### Custom Macros: dbt run-operation

Run custom macros:

```bash
dbt run-operation crazy_macro
```

### Clone Production Environment

Clone your production environment faster than you can say "duplicate":

```bash
dbt clone --state path/to/artifacts
```


# Methods

Learn how to use dbt selector methods to filter resources based on properties like tags, sources, paths, and configurations, improving the precision of dbt runs and tests.

Selector methods allow you to filter resources based on specific properties using the `method:value` syntax. While it's advisable to explicitly denote the method, you can omit it, and the default will be one of `path`, `file`, or `fqn`.

Most selector methods below support unix-style wildcards:

| Wildcard | Description                                                   | Example                                   |
| -------- | ------------------------------------------------------------- | ----------------------------------------- |
| \*       | matches any number of characters (including none)             | `dbt list --select "*.`*`folder_name.*"`* |
| ?        | matches any single character                                  | `dbt list --select "model_?.sql"`         |
| \[abc]   | matches one character listed in the bracket                   | `dbt list --select "model_[abc].sql"`     |
| \[a-z]   | matches one character from the specified range in the bracket | `dbt list --select "model_[a-z].sql"`     |

Below are examples of several popular selector methods:

### "tag" Method

Use the `tag:` method select models with a specified tag.

```bash
# Run all models with the 'hourly' tag
dbt run --select "tag:hourly"    
```

### "source" Method

Use the `source:` method to select models that reference a specified source.

```bash
# Runs all models that reference the fivetran source
dbt run --select "source:fivetran+"
```

### "resource\_type" method

Use the `resource_type` method to select nodes of a specific type (ex. `model`, `test`, `exposure`, etc.)

```bash
# Runs all models and tasks related to exposures
dbt run --select "resource_type:exposure"  
```

```bash
# Lists all tests in your project
dbt list --select "resource_type:test"  
```

### "path" method

Use `path` method to select models/sources defined at or under a specific path.

```bash
# Runs all models in the "models/marts" path
dbt run --select "path:models/marts"
```

```bash
# Runs a specific model, "customers.sql", in the "models/marts" path. 
dbt run --select "path:models/marts/customers.sql"
```

### "file" method

Use `file` method to select a model by filename.

```bash
# Runs the model defined in 'model_name.sql'
dbt run --select "file:model_name.sql"

# Note: Adding the file extension (e.g., ".sql") is optional.
dbt run --select "file:model_name"
```

### "fqn" method

Use 'fqn' to select nodes based off their "fully qualified name" (FQN). The default FQM format includes the dbt project name, subdirectories, and the file name

```bash
# Runs the model named 'example_model'
dbt run --select "fqn:example_model"

# Runs the model 'model_one' in 'example_model'
dbt run --select "fqn:project_name.example_model"

# Runs 'example_model' in 'package_name'
dbt run --select "fqn:package_name.example_model"

# Runs 'example_model' in 'example_path'
dbt run --select "fqn:example_path.example_model"

# Runs 'example_model' in 'example_path' within 'project_name'
dbt run --select "fqn:project_name.example_path.example_model"
```

### "package" method

Use `package` method to select models defined within the root project or an installed dbt package.

```bash
# Runs all models in the 'fivetran' package
dbt run --select "package:fivetran"

# Note: Adding "package" prefix is optional. The following commands are equivalent:
dbt run --select "fivetran"
dbt run --select "fivetran.*"
```

### "config" method

Use `config` to select models that match a specified node config.

```bash
# Runs all models that are materialized as tables
dbt run --select "config.materialized:table"

# Runs all models that are created in the 'stagins' schema
dbt run --select "config.schema:staging"

# Runs all models clustered by 'zip_code'
dbt run --select "config.cluster_by:zip_code"
```

**Note**: `config` method work for non-string values, such as: booleans, dictionary keys, values in arrays, etc.

Suppose you have a model with the following configurations:

```bash
{{ config(
  materialized = 'view',
  unique_key = ['customer_id', 'order_id'],
  grants = {'insert': ['sales_team', 'marketing_team']},
  transient = false
) }}

select ...
```

You can use `config` method to select the following:

```bash
# Lists all models materialized as views
dbt ls -s config.materialized:view

# Lists all models with 'customer_id' as a unique key
dbt ls -s config.unique_key:customer_id

# Lists all models with insert grants for the sales team
dbt ls -s config.grants.insert:sales_team

# Lists all models that are not transient
dbt ls -s config.transient:false
```

### "test\_type" method

Use `test_type` to select tests based on type (`singular` or `generic)`

```bash
# Runs all generic tests
dbt test --select "test_type:generic"

# Runs all singular tests
dbt test --select "test_type:singular"
```

### "test\_name" method

Use `test_name` method to select tests based on the name of the test defined.

```bash
# Runs all instances of the 'not null' test
dbt test --select "test_name:not null"

# Runs all instances of the 'dbt_utils.not_accepted_values' test
dbt test --select "test_name:not_accepted_values"
```

### "state" method

{% hint style="info" %}
When using the "state" method in a Bolt schedule of type **Deferred** or **Turbo-CI**, you don't need to pass the `--state path/to/project/artifacts` to your dbt command.

Paradime will point to the artifacts based on the [Bolt schedule configurations](/app-help/documentation/bolt/creating-schedules):

* Deferred schedule
* Last run type
  {% endhint %}

**Note:** State-based selection is a powerful and complex feature. Make sure to read about the [known caveats and limitations](https://docs.getdbt.com/reference/node-selection/state-comparison-caveats) of state comparison.

The state method selects nodes by comparing them against a previous version of the same project, represented by a manifest. The file path of the comparison manifest must be specified using the `--state` flag or the `DBT_STATE` environment variable.

* `state:new`: Indicates there is no node with the same `unique_id` in the comparison manifest.
* `state:modified`: Includes all new nodes and any changes to existing nodes.

```bash
# run all tests on new models + and new tests on old models
dbt test --select "state:new" --state path/to/artifacts

# run all models that have been modified
dbt run --select "state:modified" --state path/to/artifacts

# list all modified nodes (not just models)
dbt ls --select "state:modified" --state path/to/artifacts
```

Because state comparison is complex, and everyone's project is different, dbt supports subselectors that include a subset of the full `modified` criteria:

* `state:modified.body`: Changes to node body (e.g. model SQL, seed values)
* `state:modified.configs`: Changes to any node configs, excluding `database`/`schema`/`alias`
* `state:modified.relation`: Changes to `database`/`schema`/`alias` (the database representation of this node), irrespective of `target` values or `generate_x_name` macros
* `state:modified.persisted_descriptions`: Changes to relation- or column-level `description`, *if and only if* `persist_docs` is enabled at each level
* `state:modified.macros`: Changes to upstream macros (whether called directly or indirectly by another macro)
* `state:modified.contract`: Changes to a model's contract, which currently include the `name` and `data_type` of `columns`. Removing or changing the type of an existing column is considered a breaking change, and will raise an error.

Remember that `state:modified` includes *all* of the criteria above, as well as some extra resource-specific criteria, such as modifying a source's `freshness` or `quoting` rules or an exposure's `maturity` property.

There are two additional `state` selectors that complement `state:new` and `state:modified` by representing the inverse of those functions:

* `state:old` — A node with the same `unique_id` exists in the comparison manifest
* `state:unmodified` — All existing nodes with no changes

These selectors can help you shorten run times by excluding unchanged nodes. Currently, no subselectors are available at this time, but that might change as use cases evolve.

### "exposure" method

Use `exposure` method to select the parent resources of an exposure.

```bash
# tests all models that feed into the monthly_reports exposure
dbt test --select "exposure:monthly_reports" 

# Runs all upstream resources of all exposures
dbt run --select "+exposure:*"

# Lists all upstream models of all exposures
dbt ls --select "+exposure:*" --resource-type model  
```

### "metric" method

Use `metric` method to select parent resources of a metric.

```bash
# Runs all upstream resources of the monthly_qualified_leads metric
dbt run --select "+metric:monthly_qualified_leads"

# Builds all upstream models of all metrics
dbt build --select "+metric:*" --resource-type model
```

### "results" method

{% hint style="info" %}
When using the "results" method in a Bolt schedule of type **Deferred** or **Turbo-CI**, you don't need to pass the `--state path/to/project/artifacts` to your dbt command.

Paradime will point to the artifacts based on the [Bolt schedule configurations](/app-help/documentation/bolt/creating-schedules):

* Deferred schedule
* Last run type
  {% endhint %}

Use `result` method to select resources based on their results status from a previous execution.

```bash
# Runs all models that successfully ran on the previous execution of dbt run
dbt run --select "result:success" --state path/to/project/artifacts

# Runs all tests that issued warnings on the previous execution of dbt test
dbt test --select "result:warn" --state /path/to/project/artifacts

# Runs all seeds that failed on the previous execution of dbt seed
dbt seed --select "result:fail" --state /path/to/project/artifacts
```

**Note**: This method only works if a dbt command (ex. `seed`, `test`, `run`, `build`.) was performed prior.

### "source\_status" method

{% hint style="info" %}
When using the "source\_status" method in a Bolt schedule of type **Deferred** or **Turbo-CI**, you don't need to pass the `--state path/to/project/artifacts` to your dbt command.

Paradime will point to the artifacts based on the [Bolt schedule configurations](/app-help/documentation/bolt/creating-schedules):

* Deferred schedule
* Last run type
  {% endhint %}

Another element of job state is the `source_status` from a prior dbt invocation. For instance, after running `dbt source freshness`, dbt generates the `sources.json` artifact, which includes execution times and `max_loaded_at` dates for dbt sources.

The following dbt commands produce `sources.json` artifacts whose results can be referenced in subsequent dbt invocations:

* `dbt source freshness`

After running one of the above commands, you can reference the source freshness results by adding a selector to a subsequent command as follows:

```bash
# You can also set the DBT_STATE environment variable instead of the --state flag.
# must be run again to compare current to previous state
dbt source freshness 
dbt build --select "source_status:fresher+" --state path/to/prod/artifacts
```

### "group" method

Use `group` method to select models defined within a specified group.

```bash
# Runs all models that belong to the marketing group
dbt run --select "group:marketing"
```

### "access" Method

Use `access` method to select models based on their access property.

```bash
# List all public models
dbt list --select "access:public"
```

### "version" Method

Use `version` to select versioned models based on the following:

* **Version Identifier**: A specific version label or number (`old`, `prerelease`, `latest`)
* **Latest** **Version**: The most recent version of a model.

```bash
# lists versios older than the 'latest' version
dbt list --select "version:old"  

# Lists versions new than the 'latest' version. 
dbt list --select "version:prerelease"  

# lists the 'latest'version
dbt list --select "version:latest"  
```

### "semantic\_model" Method

Use `semantic_model` method to selects semantic models.

```bash
# Runs the semantic model named "sales" and all its dependencies
dbt run --select "semantic_model:sales"

# Builds the semantic model "customer_orders", as well as all upstream resources
dbt build --select "+semantic_model:customer_orders"  

# Lists all resources semantic models
dbt ls --select "semantic_model:*"  
```

### "saved\_query" method

Use `saved_query` method to selects saved queries.

```bash
# Lists all saved queries
dbt list --select "saved_query:*"                    

# Lists your saved query named "customers_queries" and all upstream resources
dbt list --select "+saved_query:customers_queries"  
```

### "unit\_test" method

Use `unit_test` method to selects dbt™️ unit tests.

```bash
# list all unit tests 
dbt list --select "unit_test:*"                        

# list your unit test named "orders_with_zero_items" and all upstream resources
dbt list --select "+unit_test:orders_with_zero_items"  
```


# Selector Methods

Explore dbt selector methods to filter resources based on properties like tags, sources, and configurations, with wildcard support for precise targeting during data transformations.

Selector methods in dbt allow you to filter resources based on specific properties using the `method:value` syntax. This gives you the power to target exactly what you need during your data transformations.

<figure><img src="/files/f6tEc7NHdheyYLCBAoGR" alt=""><figcaption></figcaption></figure>

### Wildcard Magic

Most selector methods support Unix-style wildcards, which can help you target a broader set of resources:

* `*`: Matches any number of characters (including none)
* `?`: Matches any single character
* `[abc]`: Matches one character listed in the bracket
* `[a-z]`: Matches one character from the specified range in the bracket

For example:

```bash
dbt list --select "*.folder_name.*"
dbt list --select "model_[a-z].sql"
```

### The Selectors

#### Tag Selector

Use `tag:` to select models with a specific tag.

```bash
dbt run --select "tag:hourly"
```

#### Source Selector

Use `source:` to select models that reference a specified source.

```bash
dbt run --select "source:fivetran+"
```

#### Resource Type Selector

Use `resource_type:` to select nodes of a specific type (e.g., model, test, exposure).

```bash
dbt run --select "resource_type:exposure"
dbt list --select "resource_type:test"
```

#### Path Selector

Use `path:` to select models/sources defined at or under a specific path.

```bash
dbt run --select "path:models/marts"
dbt run --select "path:models/marts/customers.sql"
```

#### File Selector

Use `file:` to select a model by filename.

```bash
dbt run --select "file:model_name.sql"
```

#### FQN (Fully Qualified Name) Selector

Use `fqn:` to select nodes based on their fully qualified name.

```bash
dbt run --select "fqn:example_model"
dbt run --select "fqn:project_name.example_path.example_model"
```

#### Package Selector

Use `package:` to select models defined within the root project or an installed dbt package.

```bash
dbt run --select "package:fivetran"
```

#### Config Selector

Use `config:` to select models that match a specified node config.

```bash
dbt run --select "config.materialized:table"
dbt run --select "config.cluster_by:zip_code"
```

#### Test Type Selector

Use `test_type:` to select tests based on type (generic, singular, unit, data).

```bash
dbt test --select "test_type:generic"
dbt test --select "test_type:singular"
```

#### Test Name Selector

Use `test_name:` to select tests based on the name of the test defined.

```bash
dbt test --select "test_name:not_null"
```

#### State Selector

Use `state:` to select nodes by comparing them against a previous version of the project.

```bash
dbt test --select "state:new" --state path/to/artifacts
dbt run --select "state:modified" --state path/to/artifacts
```

#### Exposure Selector

Use `exposure:` to select the parent resources of an exposure.

```bash
dbt test --select "exposure:monthly_reports"
dbt run --select "+exposure:*"
```

#### Metric Selector

Use `metric:` to select parent resources of a metric.

```bash
dbt run --select "+metric:monthly_qualified_leads"
```

#### Results Selector

Use `result:` to select resources based on their results status from a previous execution.

```bash
dbt run --select "result:success" --state path/to/project/artifacts
dbt test --select "result:warn" --state /path/to/project/artifacts
```

#### Source Status Selector

Use `source_status:` to select based on the freshness of sources.

```bash
dbt source freshness
dbt build --select "source_status:fresher+" --state path/to/prod/artifacts
```

#### Group Selector

Use `group:` to select models defined within a specified group.

```bash
dbt run --select "group:marketing"
```

#### Access Selector

Use `access:` to select models based on their access property.

```bash
dbt list --select "access:public"
```

#### Version Selector

Use `version:` to select versioned models.

```bash
dbt list --select "version:old"
dbt list --select "version:latest"
```

#### Semantic Model Selector

Use `semantic_model:` to select semantic models.

```bash
dbt ls --select "semantic_model:sales"
```

#### Saved Query Selector

Use `saved_query:` to select saved queries.

```bash
dbt list --select "saved_query:*"
```

#### Unit Test Selector

Use `unit_test:` to select dbt unit tests.

```bash
dbt list --select "unit_test:*"
```

### Pro Tips

* Combine selectors for laser-focused selection:

```bash
dbt run --select "tag:nightly,config.materialized:table"
```

* Use graph operators like `+` and `@` for complex selections:

```bash
dbt run --select "source:raw_data+,@tag:critical"
```

* Exclude models using the `--exclude` flag:

```bash
dbt run --select "path:models/mart" --exclude "tag:deprecated"
```

* If you omit the method, dbt will default to one of `path`, `file`, or `fqn`.

Remember, the more targeted your selector methods, the more precisely you can execute your dbt transformations.


# Graph Operators

Learn how to use dbt graph operators with the --select flag to navigate your project's execution graph, targeting specific models and their dependencies for precise execution in data pipelines.

Graph operators in dbt are special syntax used with the `--select` flag to target specific parts of your project's execution graph (DAG). This graph represents the dependencies between your models, with each model as a node and the dependencies as edges.

The graph operators allow you to navigate this execution graph and select subsets of your project's resources. They're like secret codes to tell dbt exactly which models you want to work with, whether it's running, testing, or listing them.

<figure><img src="/files/f6tEc7NHdheyYLCBAoGR" alt=""><figcaption></figcaption></figure>

### Operators

#### Wildcard Operator (`*`)

Run all models in a schema.

```bash
dbt run --select my_schema.*
```

#### Path Operator

No special character needed, just use the path.

```bash
dbt run --select models/staging
```

#### Parent/Child Operator (`+`)

The plus before a model name selects the model and its parents. The plus after a model name selects the model and its children.

```bash
dbt run --select +final_model
dbt run --select parent_model+
```

#### Exclusion Operator (`@`)

Select parents or children, without the original model.

```bash
dbt run --select @model_name
dbt run --select model_name@
```

#### Selection Operator (`,`)

Run multiple models.

```bash
dbt run --select model1,model2,model3
```

#### Intersection Operator (`,`)

Get the intersection of multiple selectors.

```bash
dbt run --select tag:nightly,staging.*
```

### Pro Tips

* Combine operators for laser-focused selection:

```bash
dbt run --select tag:nightly,+final_model
```

* Use `dbt ls` to preview your selection:

```bash
dbt ls --select tag:nightly,+final_model
```

* Refresh two generations of parents and all children of critical models:

```bash
dbt run --select +2tag:critical+
```

* Test everything related to final reports except the reports themselves:

```bash
dbt test --select @tag:final_report@
```

* Remember, the order of operators matters, as dbt processes them from left to right.

With these graph operators, you can create powerful, precise dbt commands to execute exactly the models you need in your data transformation pipelines.


# Paradime fundamentals


# Global Search


# Paradime Apps Navigation

Paradime Apps Navigation: Navigate and utilize Paradime applications for dbt™ projects. Enhance workflow efficiency.

Using the Paradime Global search enable to quickly move across apps, open settings, add users and preview bolt schedules runs status from anywhere in Paradime.

You can also quickly search for Paradime help docs on the fly.

To activate Global Search, click on the input on the top of your screen, or use the keyboard shortcut `cmd+K` or `ctrl+K`.

{% @arcade/embed url="<https://app.arcade.software/share/TJtCAEuaZu8rRBNnEj6M>" flowId="TJtCAEuaZu8rRBNnEj6M" %}


# Invite users to your workspace

Use Paradime Global Search to invite users directly without accessing team settings. Simply search for 'Add New User' using cmd+K or ctrl+K to send invites via email or Slack.

Using the Paradime Global search you can invite users to Paradime without having to navigate to the team settings screen.

To activate Global Search, click on the input on the top of your screen, or use the keyboard shortcut `cmd+K` or `ctrl+K`

Then search for `Add New User` to launch the popup where you can invite users to your workspace vie email or Slack.

{% @arcade/embed url="<https://app.arcade.software/share/G0CMl3Z2xoBKQItbJIaa>" flowId="G0CMl3Z2xoBKQItbJIaa" %}

**ℹ️ Check other articles on user management in Paradime:**

{% content-ref url="/pages/j9bsje7ScFQxDdGtu1IB" %}
[Users](/app-help/documentation/settings/users)
{% endcontent-ref %}


# Search and preview Bolt schedules status

Paradime Bolt Search: Search and preview the status of Bolt schedules for dbt™ projects. Ensure efficient schedule management.

Using the Paradime Global search you can search from anywhere in Paradime your Bolt schedules and preview the last status of each schedule run.

To activate Global Search, click on the input on the top of your screen, or use the keyboard shortcut `cmd+K` or `ctrl+K`.

Then type Bolt schedules or a schedule name to see the last run status. Click on one of the results to open the schedule run details and view run logs.

{% @arcade/embed url="<https://app.arcade.software/share/dcryVoFDcTBmmpOoaKlX>" flowId="dcryVoFDcTBmmpOoaKlX" %}

{% hint style="info" %}
**Check other** [**Bolt Schedule**](/app-help/documentation/bolt) **resources.**
{% endhint %}




---

[Next Page](/app-help/llms-full.txt/1)

