# What is Select Star?

Discover and understand your data effortlessly with Select Star—an intelligent data discovery platform that simplifies data exploration and documentation.

[Select Star](https://www.selectstar.com) is an intelligent data discovery platform that automatically analyzes and documents your data.

Many data scientists and business analysts spend too much time looking for the right data, often spending excessive amounts of time having to ask other people to find it.

Select Star provides an easy to use data portal that everyone can use to find and understand data.

## High Level Architecture

So, how does Select Star work?

![](/files/-MiX8eiZc-OxKf9QQGmR)

Select Star connects directly to data warehouses and BI tools to fetch metadata, query history, and activity logs from each data source.

Once the metadata and query history enters Select Star, it goes into a unified metadata store via a query parser. The query parser analyzes what data exists, where it exists, and how it's being used.

Based on that query history, popularity is automatically calculated based on what's being used, what's not being used, and who's using what. This popularity shows on every data object in Select Star.

By default, all columns in a table, all tables in a database, all dashboards in a BI tool will be ordered by popularity in Select Star so you can always see what's being used most first.

<div><figure><img src="/files/8HH8NoJ1pYPT8W00QE9q" alt=""><figcaption><p>Table overview page</p></figcaption></figure> <figure><img src="/files/pDR0fePoQblGZnqB7Yxw" alt=""><figcaption><p>Table page showing all columns ordered by popularity</p></figcaption></figure></div>

Each table and dashboard in Select Star then shows a data lineage model, to easily see from one place what the data dependencies are, where the data came from, and where it's being used.

<figure><img src="/files/lLQOYFAUulGjkHRE15IF" alt=""><figcaption></figcaption></figure>

Using Select Star, customers can build metrics, definitions, and workflows so they can organize and describe their data across all their data tools from one central location.


# Getting Started

This section provides resources to start using Select Star quickly and effectively, including tutorials, guides, and tips. Get up and running with ease!

The pages in this section will help you get up and running with Select Star quickly.

If you are setting up your organization for the first time, you can watch this video to see how to get through the onboarding flow.

{% embed url="<https://www.loom.com/share/458429a106804d7a90676162b921a179?hideEmbedTopBar=true>" %}

The steps in this guide are additional steps you can take at any point as an [Admin](/user-management/user-roles) to enable your organization to start documenting, organizing, and discovering your data with Select Star.

Connect your first data source to Select Star.

{% content-ref url="/pages/-MgSRO5bdErGcopH4VWl" %}
[1. Data Source Setup](/getting-started/data-source-setup)
{% endcontent-ref %}

Identify accounts used for maintenance or by other tools so your popularity will be more accurately reflected.

{% content-ref url="/pages/-MgSUyapCRiQ54fwu\_JL" %}
[2. Mark Service Accounts](/getting-started/mark-service-accounts)
{% endcontent-ref %}

Hide test tables and dashboards from view in Select Star so you can focus on the most relevant data to your organization.

{% content-ref url="/pages/-MgSV-Ulp6L\_JoSDaTd-" %}
[3. Hide Unwanted Datasets](/getting-started/clean-up)
{% endcontent-ref %}

Invite data team members from your organization as the main point of contact for each dataset.

{% content-ref url="/pages/-MgSV383sS7PKil8DaIR" %}
[4. Invite Owners](/getting-started/invite-owners)
{% endcontent-ref %}

Add documentation to your datasets so everyone can easily find the data they need.

{% content-ref url="/pages/-MgSV57A-SshWUAIDE\_d" %}
[5. Add Documentation](/getting-started/add-documentation-gs)
{% endcontent-ref %}

After you've done all that, you can check out Next Steps.

{% content-ref url="/pages/-MiYh6pOlzBiI8sao8kA" %}
[Next Steps](/getting-started/next-steps)
{% endcontent-ref %}


# 1. Data Source Setup

In this guide, we will walk you through how to connect your data to Select Star through our integrations and star utilizing your metadata through lineage generation.

The first thing you need to do when setting up Select Star is to connect a data source.

Visit the [Integrations](/integrations) page to see a list of integrations. Find the page for your data source to see specific setup instructions.

{% content-ref url="/pages/-MgSVP4u2TS\_45Ykm4-n" %}
[Integrations](/integrations)
{% endcontent-ref %}

The steps for granting Select Star access to your metadata are slightly different for each source, but you will need admin access to any source you want to connect.

Add your data source to Select Star by going to **Settings.**

![](/files/WDzYRlbf6RS6MliLAZzZ)

Click **Data** in the sidebar, then **+ Add** on the page.

<div data-full-width="true"><figure><img src="/files/qLJ8L6av956RuUrEGeIW" alt=""><figcaption></figcaption></figure></div>

Clicking **+ Add** will open a modal with a dropdown where you can select your data source.

{% hint style="warning" %}
If you have a data warehouse, it is strongly recommended that you add the data warehouse connection **first**, before you add BI tools and other integrations.
{% endhint %}

Once you connect your source, [lineage](/features/lineage) will be automatically generated for each data asset across your data sources, as well as the popularity of any tables, columns, and dashboards. \*\*\*\*


# 2. Mark Service Accounts

Discover how to mark and manage ETL users as service accounts with this comprehensive guide from Select Star. Learn the importance of service accounts and how to properly categorize them for optimal d

To make your popularity calculations more accurate, mark your ETL users as Service Accounts.

To mark Service Accounts, go to **Settings** > **Data** > and select your data source from the list in the left sidebar.

Check the **Service Account** box :ballot\_box\_with\_check: next to Service Account users and their activity will automatically be weighted less in popularity calculations.

{% hint style="info" %}
Click :mag:in the header or use the keyboard shortcut `Ctrl/Cmd+F` to filter users.
{% endhint %}

<figure><img src="/files/GnhxcfC71T9uVuQGVigQ" alt=""><figcaption></figcaption></figure>


# 3. Hide Unwanted Datasets

In Select Star, you can hide datasets created in personal spaces, that are used for testing purposes or that aren't relevant to the entire company. Learn how to here.

Hide datasets created in personal spaces, or for testing purposes which aren't relevant to the entire company.

Click the name of your database in the **left sidebar** to see all the schemas and tables in your data warehouse.

<figure><img src="/files/Cy2QvIsLiUlxSvt7bPkT" alt=""><figcaption></figcaption></figure>

Use `Cmd/Ctrl+F` to open a filter where you can search for specific datasets to remove, or use the filters in the right sidebar to show items with the lowest popularity.

This can help identify un-used datasets or datasets only used for testing purposes.

<div data-full-width="true"><figure><img src="/files/l8Ptf4BKfyFhfGV1Bgn2" alt=""><figcaption></figcaption></figure> <figure><img src="/files/uYdyr5xoyZ7xneO5Lkm5" alt=""><figcaption></figcaption></figure></div>

<figure><img src="/files/QmAV0IWhrK3wZgGhVXGQ" alt=""><figcaption></figcaption></figure>

Select the datasets you'd like to hide (Use `Shift` on your keyboard to select consecutive items) and click the **Hide** button.

You'll be prompted to confirm before deleting.

<figure><img src="/files/P0eLEqDtsMj09qMl01Y0" alt=""><figcaption></figcaption></figure>

You'll be able to add the data back if you need to later. See this page on [managing data sources](/data-source-management/manage-data-sources) for more information.


# 4. Invite Owners

If you're looking to improve your Select Star organization, inviting users is a great way to collaborate with others and get more done. Here's how to invite users to your Select Star organization.

Invite a few other people from your organization to help you start organizing your data.

The best candidates are usually data subject matter experts, whether they are technically-oriented or business-oriented.

These people will be the best-suited to document and tag your data as the ones who work with it most closely and know how it is constructed and used.

Invite new users by going to **Settings** > **Users**, then click **+ Invite User**.

Enter the user’s email, then select a **Team** and a [**Role**](/user-management/user-roles) for them. Your subject matter experts will most likely be either Data Managers or Admins.

<div data-full-width="true"><figure><img src="/files/Jc7U0vmzq8ZV5KcRM1Cw" alt=""><figcaption></figcaption></figure></div>

The user will receive an email inviting them to create an account on your Select Star instance. \*\*\*\*

{% hint style="info" %}
Select Star can help you bulk assign owners based on usage or who created the table. Please reach out to your CSM to request this service. Note: users must be invited, but not required to have activated their account to be set as an owner via this method. This is in order to facilitate much quicker onboarding for our customers.
{% endhint %}

Learn more about inviting users by clicking the link below.

{% content-ref url="/pages/-MgC87\_dBEELjEv-V1YV" %}
[Invite Users](/user-management/invite-users)
{% endcontent-ref %}


# 5. Add Documentation

Learn how to add searchable documentation to Select Star and help people within your organization understand your data better, fostering collaboration and better decision-making.

By adding searchable documentation to Select Star, you can help people from all different technical levels in your organization understand existing data and what it means.

You can easily [add documentation](/data-management/add-documentation) to tables and dashboards with an **Edit Description** button. Document columns by clicking in the description field.

<div data-full-width="true"><figure><img src="/files/F85yXkVaMaatlNPDHz1m" alt=""><figcaption></figcaption></figure></div>

Database pages and [Tag](/data-management/tag-management#creating-tags) pages can have descriptions as well, to give an overview of what kind of information is available.

<div data-full-width="true"><figure><img src="/files/D1VPYmzL7HTaVj4UbWxc" alt=""><figcaption></figcaption></figure></div>

You can adjust your filters in the right sidebar to show tables with higher popularity scores, but no description, to identify the best places to start.

Quickly upload large amounts of metadata in `.csv` format by going to **Settings**. Click **Metadata Upload** under the **Data Manager** section.

![](/files/ye9TWFB2GfU6GnbU67Xy)

Learn more about adding documentation by clicking the link below.

{% content-ref url="/pages/-MgSXZ2I47\_zHmehQZ3u" %}
[Add Documentation](/data-management/add-documentation)
{% endcontent-ref %}


# Next Steps

Once you have completed our easy, 5 step set up process, you are ready to start utilizing the platform to optimize your data management and analytics processes.

After you've completed these five steps, your organization is set up for others to start joining Select Star.

If you have other data tools to set up, visit the Integrations pages.

{% content-ref url="/pages/-MgSVP4u2TS\_45Ykm4-n" %}
[Integrations](/integrations)
{% endcontent-ref %}

See detailed information about specific Select Star Features, or explore Data Discovery workflows.

{% content-ref url="/pages/-MkwYaJhaNX6V2yn2mpJ" %}
[Features](/features)
{% endcontent-ref %}

{% content-ref url="/pages/-MkwZCdQadfrewbgCTsI" %}
[Data Discovery](/data-discovery)
{% endcontent-ref %}

Data Managers can learn how to document and organize data from the pages on Data Management.

{% content-ref url="/pages/-MkwZOoMZQzL0gKrqAqf" %}
[Data Management](/data-management)
{% endcontent-ref %}

Detailed information about using each of Select Star's integrations is available under Learning Data.

{% content-ref url="/pages/-MkwZXOd2wGWE0sL-ObJ" %}
[Learning Data](/learning-data)
{% endcontent-ref %}

Admins can configure data sources and users by consulting the Data Source Management and User Management sections.

{% content-ref url="/pages/-MkwZjAVQrHMQTuMU5MX" %}
[Data Source Management](/data-source-management)
{% endcontent-ref %}

{% content-ref url="/pages/-MkwZqplj0-Vb0fNpmVX" %}
[User Management](/user-management)
{% endcontent-ref %}

Developers can check out our Select Star API for setup instructions and documentation.

{% content-ref url="/pages/-Mkw\_00H3MfUYEoEtgaK" %}
[Select Star API](/select-star-api)
{% endcontent-ref %}


# Integrations

Explore this page to view a comprehensive list of all Select Star integrations and the current features they support. Stay up-to-date with new integrations and features as they become available.

This page shows a list of integrations and the features they currently support.

* [Cloud Data Warehouses](#cloud-data-warehouses)
* [BI / Data Visualization Tools](#bi-data-visualization-tools)
* [ETL Tools](#etl-and-other-tools)
* [Other Platform Apps](#other-platform-apps) (i.e., Slack)
* [SSO](#sso) (Okta, Azure AD, etc)

{% hint style="info" %}
**IP Whitelisting**

For any instance protected by a firewall, you must whitelist the following two IP addresses to connect Select Star.
{% endhint %}

```
3.23.108.85
3.20.56.105
```

For instances operating within a non-publicly accessible environment, such as an AWS VPC, please refer to our guide on [Integrating Private Network Data Sources](/integrations/private-network) for detailed instructions and best practices.

If you are setting up your organization for the first time and adding data sources through the onboarding flow, you can watch this video for more information.

{% embed url="<https://www.loom.com/share/458429a106804d7a90676162b921a179?hideEmbedTopBar=true>" %}

## Cloud Data Warehouses

| Integration                                                                         | Metadata             | Popularity           | Lineage              |
| ----------------------------------------------------------------------------------- | -------------------- | -------------------- | -------------------- |
| ![](/files/-MlC2qVQfDjFCfCzuQ4y) [Snowflake](/integrations/snowflake)               | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/Y4SIcl4zPw1OEZGoR6jS) [Databricks](/integrations/databricks)             | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/-MlC1tWt8cUoU0OOkt0e) [BigQuery](/integrations/bigquery)                 | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/-MlC2x0W54sWAlkhyf_x) [Redshift](/integrations/redshift)                 | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/-MlC2ztjyQgNjWHj6Aom) [PostgreSQL](/integrations/postgres)               | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/eMXPBEB51O3xqDbomw4n) [Microsoft SQL server (beta)](/integrations/mssql) | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/sncHVVoQCIrQdbp5RdNb) [MySQL (beta)](/integrations/mysql)                | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/T8QDcLIH4MIMVVACOCgR) [Oracle (beta)](/integrations/oracle)              | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/4BfvQUj2vBYQSwLkf81N) [Salesforce (beta)](/integrations/salesforce)      | :white\_check\_mark: |                      | :white\_check\_mark: |
| ![](/files/fPtPMCSHtosgnIO1SEy2) [DB2 (beta)](/integrations/db2)                    | :white\_check\_mark: |                      |                      |

## BI / Data Visualization Tools

| Integration                                                                                            | Metadata             | Popularity           | Lineage              |
| ------------------------------------------------------------------------------------------------------ | -------------------- | -------------------- | -------------------- |
| ![](/files/-MlC32WSH2clclFWpRcx) [Tableau](/integrations/tableau-server)                               | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/06pPhq25Oly5OdAsbtph) [Microsoft Power BI](/integrations/powerbi)                           | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/-MlC37tW0rtgtp0ST0Jf) [Looker](/integrations/looker)                                        | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/-MlC3AEriksETykLBxx9) [Mode](/integrations/mode)                                            | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/SkAohESOOMPH60pRjm0g) [Sigma](https://docs.selectstar.com/integrations/sigma-computing)     | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/63JE2azelKHyJ1CjIet3) [Sisense / Periscope](/integrations/periscope)                        | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/WAoZorFMcqQWkmDb4mgC) [Metabase](/integrations/metabase)                                    | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/nip9DwwsM69okyywpmOD) [Looker Studio](https://docs.selectstar.com/integrations/data-studio) | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/jAmY1G4TRjRK4y7wn0yI) [Thoughtspot](/integrations/thoughtspot)                              | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/FRtCi1IumHPJGU800KoE) [Quicksight](/integrations/quicksight)                                | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/FFuo8w3iXzq0F1hMLS3x) [Hex](/integrations/hex)                                              | :white\_check\_mark: | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/4BfvQUj2vBYQSwLkf81N) [Salesforce CRMA (beta)](/integrations/salesforce-analytics)          | :white\_check\_mark: |                      | :white\_check\_mark: |

## ETL & Other Tools

| Integration                                                                     | Metadata             | Lineage              |
| ------------------------------------------------------------------------------- | -------------------- | -------------------- |
| ![](/files/-MlC3D7uC9hH1_j2wLlo) [dbt](/integrations/dbt)                       | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/3JVEzzL10Smm5YORHzqZ) [AWS Glue](/integrations/aws-glue)             | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/21zDcHVNWzJfdlUfKpQx) [Apache Airflow](/integrations/apache-airflow) | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/XdNjaZqWfC2OTKBdwqSG) [OpenLineage](/integrations/openlineage)       | :white\_check\_mark: | :white\_check\_mark: |
| ![](/files/eR4NXf9lu86MwoPJXJT0) [Fivetran](/integrations/fivetran)             |                      | :white\_check\_mark: |

## Other Platform Apps

| Tool                                                          | Search               | Notifications        |
| ------------------------------------------------------------- | -------------------- | -------------------- |
| ![](/files/MG0nwXcwoBH6lM6hPIsj) [Slack](/integrations/slack) | :white\_check\_mark: | :white\_check\_mark: |

## SSO

Select Star uses [WorkOS](https://workos.com) to connect your Identity Provider to Select Star. The following integrations are supported:

* [AD FS SAML](https://workos.com/docs/integrations/adfs-saml)
* [ADP OIDC](https://workos.com/docs/integrations/adp-oidc)
* [Auth0 SAML](https://workos.com/docs/integrations/auth0-saml)
* [Azure AD SAML](https://workos.com/docs/integrations/azure-ad-saml)
* [CAS SAML](https://workos.com/docs/integrations/cas-saml)
* [ClassLink SAML](https://workos.com/docs/integrations/classlink-saml)
* [Cloudflare SAML](https://workos.com/docs/integrations/cloudflare-saml)
* [CyberArk SAML](https://workos.com/docs/integrations/cyberark-saml)
* [Duo SAML](https://workos.com/docs/integrations/duo-saml)
* [Generic SAML](https://workos.com/docs/integrations/generic-saml)
* [Google OAuth](https://workos.com/docs/integrations/g-suite-oauth)
* [Google SAML](https://workos.com/docs/integrations/google-saml)
* [JumpCloud SAML](https://workos.com/docs/integrations/jumpcloud-saml)
* [Keycloak SAML](https://workos.com/docs/integrations/keycloak-saml)
* [LastPass SAML](https://workos.com/docs/integrations/last-pass-saml)
* [Microsoft OAuth](https://workos.com/docs/integrations/microsoft-oauth)
* [miniOrange SAML](https://workos.com/docs/integrations/mini-orange-saml)
* [NetIQ SAML](https://workos.com/docs/integrations/net-iq-saml)
* [Okta SAML](https://workos.com/docs/integrations/okta-saml)
* [OneLogin SAML](https://workos.com/docs/integrations/onelogin-saml)
* [OpenID Connect](https://workos.com/docs/integrations/oidc)
* [Oracle SAML](https://workos.com/docs/integrations/oracle-saml)
* [PingFederate SAML](https://workos.com/docs/integrations/ping-federate-saml)
* [PingOne SAML](https://workos.com/docs/integrations/ping-one-saml)
* [Salesforce SAML](https://workos.com/docs/integrations/salesforce-saml)
* [Shibboleth Unsolicited SAML](https://workos.com/docs/integrations/shibboleth)
* [Shibboleth Generic SAML](https://workos.com/docs/integrations/shibboleth-generic-saml)
* [SimpleSAMLphp SAML](https://workos.com/docs/integrations/simple-saml-php-saml)
* [VMWare SAML](https://workos.com/docs/integrations/vmware-saml)

{% hint style="info" %}
You may need to create a new application with the Select Star logo as part of your setup. You can find high resolution logo image [here](https://drive.google.com/file/d/1ZKmkdxw7LwwGZe06Nf6BlcBQD1f762FG/view?usp=sharing).
{% endhint %}

Check out the [SAML SSO](/user-management/sso) page to learn more about using SSO in Select Star.


# Snowflake

## Supported Authentication Methods

{% content-ref url="/pages/EFXCWEmFvRcypmyZZcGl" %}
[Using Key Pair Authentication](/integrations/snowflake/key-pair)
{% endcontent-ref %}

{% content-ref url="/pages/uNiTkrmZYLxPfqo5SJ5D" %}
[Using Password Authentication](/integrations/snowflake/password)
{% endcontent-ref %}


# Using Key Pair Authentication

## Before you start

To connect Snowflake to Select Star, you will need:

* Admin access to your Snowflake instance via the `ACCOUNTADMIN` role.

Complete the following steps to enable metadata, lineage, and popularity for your Snowflake data in Select Star.

1. [Get Public Key from Select Star](#id-1.-get-public-key-from-select-star)
2. [Create a Select Star role and user in Snowflake](#id-2.-create-select-star-role-and-user-in-snowflake)
3. [Grant optional permissions](#id-3.-grant-optional-permissions)
4. [Connect Snowflake to Select Star](#id-4.-connect-snowflake-to-select-star)
5. [Choose databases and schemas](#id-5.-choose-databases-and-schemas)

***

## 1. Get Public Key from Select Star

1. **Go to Data Sources in Select Star**
   * Navigate to **Settings > Data**.
   * Click **+ Add** to create a new data source.

![](/files/IP5pzA0NjiBfKRqyS8wu)

2. **Select Snowflake as the Source Type**
   * In the **Source Type** dropdown, choose **Snowflake**.

![](/files/TO0ht1BqMGHKeJvypl9s)

3. **Fill in the Connection Details** Provide the following information:
   * **Display Name:**
     * Default is `Snowflake`. You can customize it if desired.
   * **Account:**
     * Your Snowflake account name (the part before `.snowflakecomputing.com` in your Snowflake URL).
   * **Role:**
     * The role you will grant to the service account user.
   * **Warehouse:**
     * The name of the data warehouse you will give us access to.
4. **Choose Key Pair Authentication**
   * Select **Key Pair** as the authentication method.

![](/files/DAwNdnrnjpvRdN7loZxt)

5. **Copy the Public Key**
   * Select Star generates a pair of keys: a **public key** and a **private key**.
   * The **public key** will be displayed on the screen along with a copy button. Use this button to copy the public key.
   * This public key is required for creating a user in Snowflake.
   * You can leave the form open, as you’ll return to it to complete the connection details after creating the Snowflake user.

## 2. Create Select Star Role and User in Snowflake

Use the `ACCOUNTADMIN` role and run the following SQL in your Snowflake instance. Replace `<PUBLIC_KEY>` with the copied public key from Select Star.

```sql
-- Required for basic metadata & query history access
CREATE ROLE selectstar_role;
GRANT IMPORTED PRIVILEGES ON DATABASE snowflake TO ROLE selectstar_role;
CREATE USER selectstar
    DEFAULT_ROLE = 'selectstar_role'
    TYPE = 'SERVICE'
    RSA_PUBLIC_KEY = '<PUBLIC_KEY>';  -- Replace with the copied public key
GRANT ROLE selectstar_role TO USER selectstar;
GRANT USAGE ON WAREHOUSE MED TO ROLE selectstar_role;
```

{% hint style="info" %}
These are the minimum permissions required for Select Star to collect basic metadata and query history. Query history is also used to generate lineage.
{% endhint %}

## 2a. Updating an Existing User with a Key Pair

If you already have a Snowflake user for Select Star and just need to add or update the key pair, use the following command:

```sql
ALTER USER selectstar SET RSA_PUBLIC_KEY = '<PUBLIC_KEY>';
```

* Replace `<PUBLIC_KEY>` with the public key generated by Select Star.

This will update the user's key pair without needing to recreate the user.

## 3. Grant optional permissions

To enable Select Star’s **Preview** feature and access additional metadata—such as **Primary Keys (PK)** and **Foreign Keys (FK)** —you’ll need to grant the following permissions.

Using the `ACCOUNTADMIN` role, execute the following SQL for each database you want to ingest (example uses `DWH` as the database name):

```sql
use role ACCOUNTADMIN;
grant usage on database DWH to role selectstar_role;
grant usage on all schemas in database DWH to role selectstar_role;
grant select on all tables in database DWH to role selectstar_role;
grant select on all views in database DWH to role selectstar_role;
grant usage on future schemas in database DWH to role selectstar_role;
grant select on future tables in database DWH to role selectstar_role;
grant select on future views in database DWH to role selectstar_role;
```

{% hint style="info" %}
If the latency of updates in `ACCOUNT_USAGE` is too high for your needs (see [Snowflake documentation on data latency](https://docs.snowflake.com/en/sql-reference/account-usage#data-latency)), you can switch to `INFORMATION_SCHEMA` instead, providing faster updates in the UI.
{% endhint %}

#### Enhanced Lineage for Dynamic Tables

To see **lineage for dynamic tables**, we recommend granting permission to **read dynamic table definitions**. Without this, lineage can only be inferred from query logs, which may not be fully reliable.

Using the `ACCOUNTADMIN` role, execute the following SQL for each database you want to ingest (example uses `DWH` as the database name):

```sql
use role ACCOUNTADMIN;
grant usage on database DWH to role selectstar_role;
grant usage on all schemas in database DWH to role selectstar_role;
grant monitor on all dynamic tables in database DWH to role selectstar_role;
grant monitor on future dynamic tables in database DWH to role selectstar_role;
```

{% hint style="success" %}
If you're granting these permissions after your Snowflake metadata has already been synced, you'll need to re-sync it.

1. Go to **Settings > Data**
2. Click on **Sync metadata** on your Snowflake Data source.
   {% endhint %}

## 4. Connect Snowflake to Select Star

1. **Complete the Setup in Select Star** Once the role has been created in Snowflake, return to Select Star to complete the setup.
2. **Fill in the Authentication Details** Provide the following information:
   * **Authentication:**
     * You already selected **Key Pair** authentication.
   * **Username:**
     * The name of the service account user you created earlier. In the example above, it is `selectstar`
3. **Test the Connection and Proceed**
   * Once all fields are filled, click **Next** to proceed.

## 5. Choose databases and schemas

After you fill in the information, you'll be asked to select the databases you'd like to load into Select Star.

{% hint style="info" %}
Select Star will not read queries or metadata or generate lineage for databases, schemas, or tables that are not loaded. Please load all data for which you expect to see lineage.

You can [change the databases and schemas](/data-source-management/manage-data-sources#configure-a-data-source) you have loaded if needed.
{% endhint %}

Select the database and click **Next**.

![](/files/nbcPV3zZdldmZTt7Bvk4)

For each database you selected, you'll be able to select the schemas.

![](/files/Jc55ZUXoWzmWILlb8S8a)

Your metadata should start loading automatically. Please allow 24-48 hours to completely generate popularity and lineage.

When the sync is complete, you'll be able to explore Snowflake in Select Star. See the link below for more information on Snowflake in Select Star.

{% hint style="info" %}
Note: To change an existing data source's configuration, go to Settings > Admin > Data Sources and click **Configure** on your Snowflake data source. Click **Back** to get to the credential screen to update credentials.
{% endhint %}

{% content-ref url="/pages/-MgSY-bV9bXpwypJdJVv" %}
[Getting Started: Snowflake](/learning-data/getting-started-snowflake)
{% endcontent-ref %}


# Using Password Authentication

## Before you start

To connect Snowflake to Select Star, you will need:

* Admin access to your Snowflake instance via the `ACCOUNTADMIN` role.

Complete the following steps to enable metadata, lineage, and popularity for your Snowflake data in Select Star.

1. [Create a Select Star role and user in Snowflake](#id-1.-create-select-star-user-in-snowflake)
2. [Grant optional permissions](#id-2.-grant-optional-permissions)
3. [Connect Snowflake to Select Star](#id-3.-connect-snowflake-to-select-star)
4. [Choose databases and schemas](#id-4.-choose-databases-and-schemas)

## 1. Create Select Star user in Snowflake

Log in to Snowflake. Using the `ACCOUNTADMIN` role, execute the following SQL:

```sql
-- Required for basic metadata & query history access
create role selectstar_role;
grant imported privileges on database snowflake to role selectstar_role;
create user selectstar password='s313ctst8r' default_role='selectstar_role' type='LEGACY_SERVICE';
grant role selectstar_role to user selectstar;
grant usage on warehouse MED to selectstar_role;
```

{% hint style="info" %}
These are the minimum permissions required for Select Star to collect basic metadata and query history. Query history is also used to generate lineage.
{% endhint %}

## 2. Grant optional permissions

To enable Select Star’s **Preview** feature and access additional metadata—such as **Primary Keys (PK)** and **Foreign Keys (FK)** —you’ll need to grant the following permissions.

Using the `ACCOUNTADMIN` role, execute the following SQL for each database you want to ingest (example uses `DWH` as the database name):

```sql
use role ACCOUNTADMIN;
grant usage on database DWH to role selectstar_role;
grant usage on all schemas in database DWH to role selectstar_role;
grant select on all tables in database DWH to role selectstar_role;
grant select on all views in database DWH to role selectstar_role;
grant usage on future schemas in database DWH to role selectstar_role;
grant select on future tables in database DWH to role selectstar_role;
grant select on future views in database DWH to role selectstar_role;
```

{% hint style="info" %}
If the latency of updates in `ACCOUNT_USAGE` is too high for your needs (see [Snowflake documentation on data latency](https://docs.snowflake.com/en/sql-reference/account-usage#data-latency)), you can switch to `INFORMATION_SCHEMA` instead, providing faster updates in the UI.
{% endhint %}

#### Enhanced Lineage for Dynamic Tables

To see **lineage for dynamic tables**, we recommend granting permission to **read dynamic table definitions**. Without this, lineage can only be inferred from query logs, which may not be fully reliable.

Using the `ACCOUNTADMIN` role, execute the following SQL for each database you want to ingest (example uses `DWH` as the database name):

```sql
use role ACCOUNTADMIN;
grant usage on database DWH to role selectstar_role;
grant usage on all schemas in database DWH to role selectstar_role;
grant monitor on all dynamic tables in database DWH to role selectstar_role;
grant monitor on future dynamic tables in database DWH to role selectstar_role;
```

{% hint style="success" %}
If you're granting these permissions after your Snowflake metadata has already been synced, you'll need to re-sync it.

1. Go to **Settings > Data**
2. Click on **Sync metadata** on your Snowflake Data source.
   {% endhint %}

## 3. Connect Snowflake to Select Star

Go to the Select Star **Settings**. Click **Data** in the sidebar, then **+ Add** to create a new Data Source.

![](/files/IP5pzA0NjiBfKRqyS8wu)

Choose **Snowflake** in the Source Type dropdown and provide the following information:

![](/files/TO0ht1BqMGHKeJvypl9s)

* **Display Name:** This value is `Snowflake` by default, but you can override it if desired.
* **Account:** The account name is the name to the left of `snowflakecomputing.com` when you log in to Snowflake.
* **Role:** The role you granted the service account user. In the example above, it is `selectstar_role`
* **Warehouse:** The name of the data warehouse you've given us access to. In the example above it is `MED`

Click **Save** and fill in the Authentication Details:

* **Authentication:** Select **Password**.
* **Username:** The name of the service account user you created. In the example above, it is `selectstar`
* **Password:** The password for the service account user you created. In the example above, it is `s313ctst8r`

![](/files/UOlmgjDuwv9UrdAnWXq4)

## 4. Choose databases and schemas

After you fill in the information, you'll be asked to select the databases you'd like to load into Select Star.

{% hint style="info" %}
Select Star will not read queries or metadata or generate lineage for databases, schemas, or tables that are not loaded. Please load all data for which you expect to see lineage.

You can [change the databases and schemas](/data-source-management/manage-data-sources#configure-a-data-source) you have loaded if needed.
{% endhint %}

Select the database and click **Next**.

![](/files/nbcPV3zZdldmZTt7Bvk4)

For each database you selected, you'll be able to select the schemas.

![](/files/Jc55ZUXoWzmWILlb8S8a)

Your metadata should start loading automatically. Please allow 24-48 hours to completely generate popularity and lineage.

When the sync is complete, you'll be able to explore Snowflake in Select Star. See the link below for more information on Snowflake in Select Star.

{% content-ref url="/pages/-MgSY-bV9bXpwypJdJVv" %}
[Getting Started: Snowflake](/learning-data/getting-started-snowflake)
{% endcontent-ref %}


# Snowflake Tag Sync

Follow these steps to enable sync of Select Star tags to Snowflake.

You have the option of syncing [Tags](/features/tags#filter) created in Select Star back to Snowflake. It will sync status and category tags for both tables and columns. Suggested tags will also be synced.

## Enable Snowflake Tag Sync in Select Star

### 1. Create a custom Tag role in Snowflake

First, you will need to create a custom role and give it access to create and apply tags.

<pre class="language-sql"><code class="lang-sql">-- Create tag_admin role
USE ROLE USERADMIN;
CREATE ROLE tag_admin;
GRANT ROLE tag_admin TO USER &#x3C;select_star_user>;


-- Enable tag creation and assignment
USE ROLE ACCOUNTADMIN;
-- &#x3C;db_name>.&#x3C;schema_name> will be where Select Star 
GRANT CREATE TAG ON SCHEMA &#x3C;db_name>.&#x3C;schema_name> TO ROLE tag_admin;
GRANT APPLY TAG ON ACCOUNT TO ROLE tag_admin;

-- Grant usage
GRANT USAGE ON DATABASE &#x3C;db_name> TO ROLE tag_admin;
<strong>GRANT USAGE ON SCHEMA &#x3C;db_name>.&#x3C;schema_name> TO ROLE tag_admin;
</strong></code></pre>

You can find more detailed instructions in Snowflake's documentation [here](https://docs.snowflake.com/en/user-guide/object-tagging.html#step-1-create-a-custom-role-and-assign-privileges).

### 2. Authorize Tag Sync in Select Star

To enable Snowflake Tag Sync in Select Star

1. Go to **Settings**.
2. Choose the **Snowflake data source** in the sidebar.
3. Click the **Snowflake Tag Sync** tab.

![](/files/pU5Zv6HDPJdgauaeGV0d)

4. Click the **Enable Tag Sync** button.

![Snowflake Tag Sync tab in Select Star Settings.](/files/WMnkQ3qSHpGkLSOAy3Bc)

5. Enter the **Role name** of the role you created in [step 1](#1.-create-a-custom-tag-role-in-snowflake) of this guide (`tag_admin` in our example).
6. Enter the **Database Name** and **Schema Name** where Select Star should create Snowflake Tags.
7. Click **Connect**.

<figure><img src="/files/rcqMcvrtujF8m7VXrgR1" alt="" width="375"><figcaption></figcaption></figure>

Select Star will be able to use this role with the credentials you added when you connected Snowflake as outlined in our [Snowflake Setup Guide](/integrations/snowflake).

You will first see a message indicating that the Sync is in progress. Once the sync is done, we'll indicate that the tags were synced successfully.

![](/files/CFBapbrnq0LjdJfWMjLO) ![](/files/z8nrPnmi4jUHc4yOqsrN)

{% hint style="info" %}
Please note: It can take a few hours for tags to update in Snowflake's `tag_references` view.
{% endhint %}

{% hint style="danger" %}
You can disable tags at any time by clicking the Disable Tags Sync button.

Disabling tags will delete all tags created by Select Star from your Snowflake instance.
{% endhint %}

## Select Star tags in Snowflake

When Snowflake Tag Sync is enabled, any Tags created in Select Star will be visible in Snowflake.

Select Star tags will have `SELECT_STAR_CATEGORY` (for Category) or `SELECT_STAR_STATUS` (for Status) prepended in the `TAG_NAME` field. The value of the tag is left empty.

![](/files/bvXcaTkSocj6OkoOC38e) ![](/files/TDEi45lNFuqAr2zLatJW)

**Example**: If you had a Sales Tag in Select Star created as a Category tag. The Snowflake tag created would be called `SELECT_STAR_CATEGORY_SALES`.

{% hint style="info" %}
Note: Currently, you can only sync tags from Select Star to Snowflake, and not the other way around. This is to prevent the tools from overwriting each other.
{% endhint %}

{% hint style="danger" %}
Tag names in Select Star must follow Snowflake's [identifier requirements](https://docs.snowflake.com/en/sql-reference/identifiers-syntax) in order to be synced to Snowflake. Spaces in tag names will be automatically replaced with an underscore.
{% endhint %}


# Snowflake Key Pair Rotation

## Introduction

This documentation provides step-by-step instructions for performing key-pair rotation in Snowflake using the Select Star API. The rotation process ensures that your authentication keys are updated securely and efficiently. All steps in this guide are executed via API requests, allowing seamless integration into automated workflows.

Here is an overview of the steps covered:

1. [Get a new public key from Select Star](#id-1.-get-a-new-public-key-from-select-star)
2. [Update the public key in Snowflake](#id-2.-update-the-public-key-in-snowflake)
3. [Update the public key of your data source in Select Star](#id-3.-update-the-public-key-of-your-data-source-in-select-star)

**Note:** The only method to rotate keys for a Snowflake integration is through API requests. Ensure that your workflows accommodate this API-based approach.

## Authentication

All API requests require an API token. Instructions to obtain and manage your API token can be found [here](/select-star-api/authentication).

Base URL: `https://api.production.selectstar.com/`

## Permissions

Role "Admin" is required to perform key-pair rotation.

## 1. Get a new public key from Select Star

This API generates a new key pair for your organization. The private key is securely stored by Select Star, while the public key is returned in the response.

**Endpoint:**

```
GET /v1/data-sources/keygen/
```

**Response:**

```json
{
  "public_key": "MIIBIjANB...."
}
```

## 2. Update the public key in Snowflake

To update the public key in Snowflake, follow the steps outlined in the [Snowflake documentation](https://docs.snowflake.com/en/user-guide/key-pair-auth#configuring-key-pair-rotation).

Assign the public key generated in the previous step to the user in Snowflake (step 1 in the Snowflake documentation).

Once the new public key is added, both the new and old keys remain valid simultaneously. This allows you to update your data source configurations in Select Star without experiencing any downtime.

**Note:** We recommend removing the old public key from the user (step 3 in the Snowflake documentation) **after** updating the data source in Select Star.

## 3. Update the public key of your data source in Select Star

This API allows users to update credentials, merging new information with existing ones. This is useful for updating public keys after rotation.

**Endpoint:**

```
PATCH /v1/data-sources/<guid>/credentials/
```

**Payload:**

```json
{
  "credentials": {
    "public_key": "MIIBIjANB....",
    // Optional other credential fields
  }
}
```


# Cortex Analyst (beta)

Cortex Analyst allows users to query Snowflake data using natural language.

1. [Permissions](#id-1.-permissions)
2. [Enable Cortex Analyst in Select Star](#id-2.-enable-cortex-analyst-in-select-star)
3. [Generate Semantic Views](#id-3.-generate-semantic-views)
4. [Leverage Semantic Views in Your Chatbot](#id-4.-leverage-semantic-views-in-your-chatbot)

## 1. Permissions

We recommend creating a **dedicated role** for Cortex usage. This helps keep ingestion and AI access patterns separate, and ensures only the necessary permissions are granted for Cortex-related use cases.

{% hint style="danger" %}
**Key-pair authentication is required.** See [Using Key Pair Authentication](/integrations/snowflake/key-pair) for setup instructions.
{% endhint %}

To use Select Star with Snowflake Cortex, your Snowflake role must be granted the following permissions:

* **USAGE** on the database and schema where semantic views will be created.
* **CREATE SEMANTIC VIEW** on the schema ([Snowflake docs](https://docs.snowflake.com/en/sql-reference/sql/create-semantic-view)).
* **SELECT** on all tables and views referenced by the semantic view.

```sql
-- Create a new role
CREATE ROLE <role_name>;

-- Grant the new role to the Select Star user
GRANT ROLE <role_name> TO USER <select_star_user>;

-- Grant usage on the database and schema where semantic views will be created
GRANT USAGE ON DATABASE <db_name> TO ROLE <role_name>;
GRANT USAGE ON SCHEMA <schema_name> TO ROLE <role_name>;

-- Grant the privilege to create semantic views in that schema
GRANT CREATE SEMANTIC VIEW ON SCHEMA <db_name>.<schema_name> TO ROLE <role_name>;

-- Grant select privileges on the underlying tables/views the semantic view will use
GRANT SELECT ON TABLE MY_DB.MY_SCHEMA.TABLE_A TO ROLE <role_name>;
GRANT SELECT ON TABLE MY_DB.MY_SCHEMA.TABLE_B TO ROLE <role_name>;
```

{% hint style="info" %}
For more details, see [Snowflake: CREATE SEMANTIC VIEW](https://docs.snowflake.com/en/sql-reference/sql/create-semantic-view) and [Key-Pair Authentication](/integrations/snowflake/key-pair).
{% endhint %}

## 2. Enable Cortex Analyst in Select Star

To get Cortex Analyst up and running:

1. Confirm that you have the **Admin** role.
2. Navigate to **Settings**.
3. In the sidebar, select your **Snowflake** data source.
4. Click the **Cortex Analyst** tab.
5. Hit the **Enable** button.

![Enable Cortext Analyst](/files/jTpQxEXgxAwOjFlSiFkJ)

You'll be prompted to provide:

* **Database Name:** This is where your semantic views will be created.
* **Schema Name:** This is the schema within that database where the semantic views will live.
* **Role:** The role that has the necessary permissions to create semantic views and access the underlying tables.

## 3. Generate Semantic Views

Once Cortex Analyst is enabled, you can generate semantic views directly from the dashboard list page (for any data source) or the table list page (for Snowflake data sources only).

![Show AI assist button](/files/YOlJPPBI8CMxf4P0RCZB)

1. Go to the **dashboards page**.
2. Click the **AI Assist** button.
3. Select **Generate Snowflake Semantic View**.

A modal will appear, asking you to provide a **Display Name** and a **Description** for your semantic views.

![Generate Semantic View](/files/W4m7899aE2nOTbJPw9da)

## 4. Leverage Semantic Views in Your Chatbot

Now that your semantic views are generated, they're automatically available to all users in your organization. They'll provide rich, relevant context for your chatbot, making interactions more informed and insightful.

![Cortext Analyst Ask AI](/files/D7prFNNJRqSiSnTWO73U)


# Databricks

Select Star supports two methods for collecting lineage data from Databricks:

* **System tables (Recommended)**: Uses Databricks system tables via SQL REST API for better performance and scalability
* **API**: Uses individual REST API calls per column

Both methods provide identical functionality and lineage coverage. The System tables method is recommended for larger deployments as it eliminates API rate limits and reduces processing time.

## Supported Hosting Solutions

{% content-ref url="/pages/QHKcU0mdrNNsW8grIpnu" %}
[Databricks on AWS](/integrations/databricks/databricks-aws)
{% endcontent-ref %}

{% content-ref url="/pages/Fxd5ShZBAX9iCH3SLaon" %}
[Databricks on Azure](/integrations/databricks/databricks-azure)
{% endcontent-ref %}

If none of the available integrations fit your needs, please feel free to contact us so that we can explore available solutions.

## Learn more

{% content-ref url="/pages/KGlgaHmYbTKOq1TVMuVR" %}
[Getting Started: Databricks](/learning-data/getting-started-databricks)
{% endcontent-ref %}


# Databricks on AWS

Learn how to connect to your Databricks data warehouse on AWS to Select Star and retrieve metadata from your datasets. This guide will provide you with the necessary steps to get started.

## **Before you start**

{% hint style="info" %}
Ensure Unity Catalog is enabled for your Databricks instance. For details, see [Getting Started with Unity catalog](https://docs.databricks.com/data-governance/unity-catalog/get-started.html).
{% endhint %}

To connect Databricks to Select Star, you will need...

* an Databricks instance on AWS. For details, see [Databricks' documentation](https://www.databricks.com/product/aws).
* Account admin permissions on the Databricks instance
* Workspace admin permissions on the Databricks instance

Complete all of the following steps to see Databricks metadata, lineage, and popularity in Select Star.

1. [Create a service principal (SelectStar) in Databricks](#id-1.-create-a-service-principal-in-databricks)
2. [Generate a Personal Access Token](#id-2.-generate-a-personal-access-token)
3. [Configure System tables lineage (Recommended)](#id-3.-configure-system-tables-lineage-recommended)
4. [Connect Databricks to Select Star](#id-4.-connect-databricks-to-select-star)
5. [Choose Catalogs and Schemas](#id-5.-choose-catalogs-and-schemas)

## **1. Create** a Service Principal **in Databricks**

#### What is a Service Principal?

A service principal is an identity that you create in Databricks for use with automated tools, jobs, and applications. Service principals give automated tools and scripts API-only access to Databricks resources, providing greater security than using users or groups. It also prevents jobs and automations from failing if a user leaves your organization or a group is modified. For details, see [Manage Service Principal](https://docs.databricks.com/administration-guide/users-groups/service-principals.html#manage-service-principals).

### **Add a service principal to your Databricks account**

Account admins can add service principals to your Databricks account using the account console or the System for Cross-domain Identity Management (SCIM) Account API.

### **Add service principals to your account using the account console**

To add a service principal to the account using the account console:

1. As an account admin, log in to the [account console](https://accounts.cloud.databricks.com/).
2. Click **User management**.
3. On the **Service principals** tab, click **Add service principal**.
4. Enter a name (**SelectStar**) for the service principal.
5. Click **Add**.

To add a service principal via REST API, see [Add service principals to your account using the SCIM (Account) API](https://docs.databricks.com/administration-guide/users-groups/service-principals.html#add-service-principals-to-your-account-using-the-scim-account-api) .

{% hint style="info" %}
💡 To use service principals, you must add them to a workspace and generate access tokens for them in the workspace.
{% endhint %}

### **Add a service principal to a workspace**

Account admins can add service principals to [identity-federated workspaces](https://docs.databricks.com/administration-guide/users-groups/index.html#assign-users-to-workspaces) using the following:

* The account console
* The Workspace Assignment API

Workspace admins can manage service principals in their workspace using the following:

* The workspace admin console (if the workspace is enabled for identity federation)
* The workspace-level SCIM (ServicePrincipals) API
* The Workspace Assignment API (if the workspace is enabled for identity federation)

### **Assign a service principal to a workspace using the account console**

To add service principals to a workspace using the account console, the workspace must be enabled for identity federation.

1. As an account admin, log in to the [account console](https://accounts.cloud.databricks.com/).
2. Click **Workspaces**.
3. On the **Permissions** tab, click **Add permissions**.
4. Search for and select the service principal **SelectStar** and assign the permission level (workspace **Admin**), and click **Save**.

To add a service principle to a workspace via admin console or REST API, see [Add a service principal to a workspace](https://docs.databricks.com/administration-guide/users-groups/service-principals.html#add-sp-workspace).

These are the minimum permissions required for Select Star to collect basic metadata and query history. Query history is also used to generate [Data Lineage](/features/lineage).

## Grant SQL and Workspace access **for a service principal**

To grant SQL Warehouse access for a service principal using the workspace admin console, the workspace must be enabled for identity federation.

1. As a workspace admin, log in to the Databricks workspace.
2. Click your username in the top bar of the Databricks workspace and select **Admin Console**.

   ![Admin Console](/files/MKoDfviD98GfxfBgp63R)
3. Click **Settings** and select **Service principals**.
4. On the **Service principals** tab, click the service principal that was create in the previous steps.
5. Select the checkbox for **Databricks SQL access** and **Workspace access**, and click **Update**.

   ![Entitlements for service principal](/files/eDivgDOzA25CMziy2wz4)

## Grant permissions to a catalog for a service principal

1. Log in to a workspace that is linked to the metastore.
2. Click **Data**.
3. Click the **catalog** that needs to be granted access to, and select **Permissions**.

   <figure><img src="/files/F3hbd3nNazoa8oEFjFCE" alt=""><figcaption><p>Catalog permissions in the Data Explorer UI</p></figcaption></figure>
4. Click **Grant**.
5. Select the user/group and grant Privilege presets to **Data Reader**, and select the checkbox for **USE CATALOG, USE SCHEMA** and **SELECT**, and click **Grant**.

![Privileges for service principal or User groups](/files/9n4uPAAMpGkxv15vNe0T)

## Grant permission to a workspace for a service principal

This step is required to show notebooks in the catalog and notebook lineage.

1. Log in to a workspace that is linked to the metastore.
2. Click **Workspace** and select top folder.
3. Click **Share** button.

   <figure><img src="/files/R9iUdxP4tLhhddhtgQTi" alt=""><figcaption><p>Folder permissions in the Workspace explore UI</p></figcaption></figure>
4. Select the user/group, then select permission "Can view", and click **Add**.

   <figure><img src="/files/KlTjiC5laAMMf2UGBeiK" alt=""><figcaption><p>Permission grant in Workspace share</p></figcaption></figure>

## **2. Generate a Personal Access Token**

To authenticate a service principal to APIs on Databricks, an administrator can create a Databricks Personal Access Tokens on behalf of the service principal.

1. Grant the [Can Use token permission](https://docs.databricks.com/administration-guide/access-control/tokens.html#control-who-can-use-or-create-tokens) to the service principal.
2. Create a Databricks personal access token on behalf of the service principal using the `POST /token-management/on-behalf-of/tokens` operation in the [token management REST API](https://docs.databricks.com/dev-tools/api/latest/token-management.html). An administrator can also list personal access tokens and delete them using the same API.

## Generate a Personal Access Token

<mark style="color:green;">`POST`</mark> `https://<deployment name>.cloud.databricks.com/api/2.0/token-management/on-behalf-of/tokens/`

When you want to use the Databricks API to generate a Personal Access token on behalf of a user or service principal, use this command.

Use the `token value` generated from this response as API key.

#### Request Body

| Name                                              | Type   | Description                                                                                                             |
| ------------------------------------------------- | ------ | ----------------------------------------------------------------------------------------------------------------------- |
| application\_id<mark style="color:red;">\*</mark> | String | UUID of the Service Principal, and can be found here - <https://accounts.cloud.databricks.com/users/serviceprincipals/> |
| comment                                           | String |                                                                                                                         |
| lifetime\_seconds                                 | String | Use value = `-1` in order for it to live indefinitely                                                                   |

{% tabs %}
{% tab title="200: OK " %}

```json

{
    "token_value": "dapia.....", #Use this value
    "token_info": {
        "token_id": "4305bc67998.........",
        "creation_time": 1671720121149,
        "expiry_time": -1,
        "comment": "Service Principal Token. API Test",
        "created_by_id": 355825636633264,
        "created_by_username": "prat@getselectstar.com",
        "owner_id": 4012126671306509
    }
}

```

{% endtab %}

{% tab title="400: Bad Request Invalid Application ID" %}

```json

{
    "error_code": "INVALID_PARAMETER_VALUE",
    "message": "edxiueirxxxxxxxx does not exist"
}

```

{% endtab %}
{% endtabs %}

For detailed, step-by-step instructions for creating access tokens for service principals, see [Service principals for Databricks automation](https://docs.databricks.com/dev-tools/service-principals.html).

## **3. Configure System tables lineage (Recommended)**

{% hint style="info" %}
💡 This section is optional but recommended. System tables lineage provides better performance and scalability by using Databricks system tables instead of individual API calls. If you skip this section, Select Star will use API lineage collection.
{% endhint %}

System tables lineage requires additional permissions beyond the basic setup. These permissions allow Select Star to query Databricks system tables that contain lineage metadata, without accessing your actual data.

### **Grant SQL Warehouse access permissions**

The service principal needs permission to use a specific SQL Warehouse for executing lineage queries.

1. In your Databricks workspace, go to **SQL Warehouses**.
2. Select the SQL Warehouse you want to use for Select Star.
3. Click the **Permissions** button.
4. Click **Add** and search for your **SelectStar** service principal.
5. Grant **Can use** permission and click **Add**.

<figure><img src="/files/wiJWUOLfKXMcf5I7W8gY" alt=""><figcaption><p>Grant CAN USE permission on SQL Warehouse</p></figcaption></figure>

{% hint style="info" %}
💡 Note the **Warehouse ID** from the SQL Warehouse details page - you'll need this when connecting to Select Star.
{% endhint %}

<figure><img src="/files/dG1MClQn0ycInsY1a4Kd" alt=""><figcaption><p>SQL Warehouse ID location</p></figcaption></figure>

### **Grant system.access schema permissions**

The service principal needs permissions to read lineage data from Databricks system tables.

1. In your Databricks workspace, go to **Catalog**.
2. Select the **system** catalog.
3. Select the **access** schema.
4. Go to the **Permissions** tab.
5. Click **Grant** and search for your **SelectStar** service principal.
6. Select **USE** and **SELECT** permissions.
7. Click **Grant**.

<figure><img src="/files/rNlPL4ja6IiAP5Yz0miR" alt=""><figcaption><p>Grant permissions on system.access schema</p></figcaption></figure>

### **Ensure SQL access entitlement**

Verify that your service principal has the SQL access entitlement enabled:

1. In your Databricks workspace, click your username and select **Admin Console**.
2. Click **Settings** and select **Service principals**.
3. Click on your **SelectStar** service principal.
4. Ensure **Databricks SQL access** is checked and click **Update** if needed.

<figure><img src="/files/powEAhsOMnjwOSMawJUT" alt=""><figcaption><p>Enable SQL access for service principal</p></figcaption></figure>

## **4. Connect Databricks to Select Star**

Go to the Select Star **Settings**. Click **Data** in the sidebar, then **+ Add** to create a new Data Source.

<figure><img src="/files/9BYfkitHCNhECWwMfX9v" alt=""><figcaption></figcaption></figure>

Choose **Databricks** in the Source Type dropdown and provide the following information:

<figure><img src="/files/GEGxuKC1kjv42prllmDR" alt=""><figcaption></figcaption></figure>

**Display Name:** This value is `Databricks` by default, but you can override it if desired.

**Workspace URL:** This is the address of the Workspace. This should include the `<deployment name>.cloud.databricks.com` . Deployment Name can be found in <https://accounts.cloud.databricks.com/workspaces>

**Access Token:** This is the **Personal access token** from Step 2, which is used to authenticate access to Databricks.

**Lineage Method:** Choose between System tables (recommended) or API lineage collection.

**SQL Warehouse ID:** Required when using System tables lineage. This is the Warehouse ID noted in Step 3. Not available for use with API lineage.

## **5. Choose Catalogs and Schemas**

After you fill in the information, you'll be asked to select the catalog you'd like to load into Select Star.

{% hint style="info" %}
💡 Select Star will not read queries or metadata or generate lineage for Catalogs, schemas, or tables that are not loaded. Please load all data for which you expect to see lineage.
{% endhint %}

You can [change the catalogs and schemas](https://docs.selectstar.com/data-source-management/manage-data-sources#configure-a-data-source) you have loaded if needed.

Select the catalogs and click **Next**.

<figure><img src="/files/RapsV8PCuymb1MAfXXAL" alt=""><figcaption></figcaption></figure>

For each catalog you selected, you'll be able to select the schemas.

<figure><img src="/files/EOScBnNVbuQhqqs2OmoU" alt=""><figcaption></figcaption></figure>

Your metadata should start loading automatically. Please allow 24-48 hours to completely generate popularity and lineage.

When the sync is complete, you'll be able to explore Databricks in Select Star.

See the link below for more information on Databricks in Select Star.

{% content-ref url="/pages/KGlgaHmYbTKOq1TVMuVR" %}
[Getting Started: Databricks](/learning-data/getting-started-databricks)
{% endcontent-ref %}


# Databricks on Azure

Learn how to connect to your Databricks data warehouse on Azure to Select Star and retrieve metadata from your datasets. This guide will provide you with the necessary steps to get started.

## **Before you start**

{% hint style="info" %}
Ensure Unity Catalog is enabled for your Databricks instance. For details, see [Getting Started with Unity Catalog](https://docs.databricks.com/data-governance/unity-catalog/get-started.html).
{% endhint %}

To connect Databricks to Select Star, you will need:

* A Databricks instance on Azure. For details, see [Databricks' documentation](https://www.databricks.com/product/azure)
* Account admin permissions on the Databricks instance
* Workspace admin permissions on the Databricks instance

Complete all of the following steps to see Databricks metadata, lineage, and popularity in Select Star:

1. [Create a service user in Databricks](#id-1.-create-a-service-user-in-databricks)
2. [Assign the service user to a workspace using the account console](#id-2.-assign-a-service-user-to-a-workspace-using-the-account-console)
3. [Grant SQL and Workspace access to the service user](#id-3.-grant-sql-and-workspace-access-for-a-service-user)
4. [Grant service user permissions to the catalog](#id-4.-grant-permissions-to-a-catalog-for-a-service-user)
5. [Grant permission to a workspace for a service user](#id-5.-grant-permission-to-a-workspace-for-a-service-user)
6. [Generate an access token](#id-6.-generate-an-access-token)
7. [Configure System tables lineage (Recommended)](#id-7.-configure-system-tables-lineage-recommended)
8. [Connect Databricks to Select Star](#id-8.-connect-databricks-to-select-star)
9. [Choose Catalogs and Schemas](#id-9.-choose-catalogs-and-schemas)

## 1. Create a Service User in Databricks

Account admins can add users to the Databricks account using the account console or the SCIM Account API. These instructions focus on using the account console approach.

To add a service user to the account using the account console:

1. As an account admin, log in to the [account console](https://accounts.azuredatabricks.net/).
2. Click **User management**.
3. On the **Users** tab, click **Add user**.
4. Enter any email, first name and last name for the service user.
5. Click **Add**.

{% hint style="info" %}
💡 To use a service user, you must be able to successfully authenticate to it. Depending on the authentication method you have configured for Databricks account (e.g. SAML), you may also need to create a service user in the corporate identity provider, such as Microsoft Entra ID.
{% endhint %}

## 2. Assign the Service User to a Workspace Using the Account Console

Account admins can add service users to [identity-federated workspaces](https://docs.databricks.com/administration-guide/users-groups/index.html#assign-users-to-workspaces) using the following:

* The account console
* The Workspace Assignment API

The following instructions focus on using the account console approach.

To add a service user to a workspace using the account console, the workspace must be enabled for identity federation.

1. As an account admin, log in to the [account console](https://accounts.cloud.databricks.com/).
2. Click **Workspaces**.
3. On the **Permissions** tab, click **Add permissions**.
4. Search for and select the service user and assign the permission level (workspace **Admin**), and click **Save**.

These are the minimum permissions required for Select Star to collect basic metadata and query history. Query history is also used to generate [Data Lineage](/features/lineage).

## 3. Grant SQL and Workspace Access to the Service User

To grant SQL Warehouse access for a service user using the workspace admin console, the workspace must be enabled for identity federation.

1. As a workspace admin, log in to the Databricks workspace.
2. Click your username in the top bar of the Databricks workspace and select **Admin Settings**.
3. Go to the **Identity and access** tab, and under **Users** click **Manage**.
4. On the **User** tab, click the service user that was create in the previous steps.
5. Select the checkbox for **Databricks SQL access** and **Workspace access**, and click **Update**.

   ![Entitlements for service user](/files/eDivgDOzA25CMziy2wz4)

## 4. Grant Service Users Permissions to the Catalog

1. As a workspace admin, log in to a workspace that is linked to the metastore.
2. Click **Catalog**.
3. Click the catalog that needs to be granted access to, and select **Permissions**.

   <figure><img src="/files/F3hbd3nNazoa8oEFjFCE" alt=""><figcaption><p>Catalog permissions in the Catalog Explorer UI</p></figcaption></figure>
4. Click **Grant**.
5. Select the user/group and grant Privilege presets to **Data Reader**, and select the checkbox for **USE CATALOG, USE SCHEMA** and **SELECT**, and click **Grant**.

![Privileges for service user or User groups](/files/9n4uPAAMpGkxv15vNe0T)

## 5. Grant permission to a workspace for a service user

This step is required to show notebooks in the catalog and notebook lineage.

1. Log in to a workspace that is linked to the metastore.
2. Click **Workspace** and select top folder.
3. Click **Share** button.

   <figure><img src="/files/R9iUdxP4tLhhddhtgQTi" alt=""><figcaption><p>Folder permissions in the Workspace explore UI</p></figcaption></figure>
4. Select the user/group, then select permission "Can view", and click **Add**.

   <figure><img src="/files/KlTjiC5laAMMf2UGBeiK" alt=""><figcaption><p>Permission grant in Workspace share</p></figcaption></figure>

## 6. Generate an Access Token

To authenticate a service user to APIs on Databricks, an administrator can create a Access Tokens.

1. As a **service user**, log in to a workspace.
2. Click your username in the top bar of the Databricks workspace and select **Admin Settings**. Ensure that the visible username is the service user you created in the previous steps.
3. Go to the **Developer** tab, and under **Access tokens** click **Manage**.
4. Click **Generate new token** and fill out the form. Once submitted, preserve access token for later use.

## **7. Configure System tables lineage (Recommended)**

{% hint style="info" %}
💡 This section is optional but recommended. System tables lineage provides better performance and scalability by using Databricks system tables instead of individual API calls. If you skip this section, Select Star will use API lineage collection.
{% endhint %}

System tables lineage requires additional permissions beyond the basic setup. These permissions allow Select Star to query Databricks system tables that contain lineage metadata, without accessing your actual data.

### **Grant SQL Warehouse access permissions**

The service user needs permission to use a specific SQL Warehouse for executing lineage queries.

1. In your Databricks workspace, go to **SQL Warehouses**.
2. Select the SQL Warehouse you want to use for Select Star.
3. Click the **Permissions** button.
4. Click **Add** and search for your service user.
5. Grant **Can use** permission and click **Add**.

<figure><img src="/files/wiJWUOLfKXMcf5I7W8gY" alt=""><figcaption><p>Grant CAN USE permission on SQL Warehouse</p></figcaption></figure>

{% hint style="info" %}
💡 Note the **Warehouse ID** from the SQL Warehouse details page - you'll need this when connecting to Select Star.
{% endhint %}

<figure><img src="/files/dG1MClQn0ycInsY1a4Kd" alt=""><figcaption><p>SQL Warehouse ID location</p></figcaption></figure>

### **Grant system.access schema permissions**

The service user needs permissions to read lineage data from Databricks system tables.

1. In your Databricks workspace, go to **Catalog**.
2. Select the **system** catalog.
3. Select the **access** schema.
4. Go to the **Permissions** tab.
5. Click **Grant** and search for your service user.
6. Select **USE** and **SELECT** permissions.
7. Click **Grant**.

<figure><img src="/files/rNlPL4ja6IiAP5Yz0miR" alt=""><figcaption><p>Grant permissions on system.access schema</p></figcaption></figure>

### **Ensure SQL access entitlement**

Verify that your service user has the SQL access entitlement enabled:

1. In your Databricks workspace, click your username and select **Admin Settings**.
2. Go to the **Identity and access** tab, and under **Users** click **Manage**.
3. Click on your service user.
4. Ensure **Databricks SQL access** is checked and click **Update** if needed.

<figure><img src="/files/BCSwWnSrFD0EUeQyUwaz" alt=""><figcaption><p>Enable SQL access for service user</p></figcaption></figure>

## **8. Connect Databricks to Select Star**

Go to the Select Star **Settings**. Click **Data** in the sidebar, then **+ Add** to create a new Data Source.

<figure><img src="/files/9BYfkitHCNhECWwMfX9v" alt=""><figcaption></figcaption></figure>

Choose **Databricks** in the Source Type dropdown and provide the following information:

<figure><img src="/files/GEGxuKC1kjv42prllmDR" alt=""><figcaption></figcaption></figure>

**Display Name:** This value is `Databricks` by default, but you can override it if desired.

**Workspace URL:** This is the address of the Workspace. This should include the `<identifier>.azuredatabricks.net`.

**Access Token:** This is the **Access token**, which is used to authenticate access to Databricks on Azure.

**Lineage Method:** Choose between System tables (recommended) or API lineage collection.

**SQL Warehouse ID:** Required when using System tables lineage. This is the Warehouse ID noted in Step 7. Not available for use with API lineage.

## **9. Choose Catalogs and Schemas**

After you fill in the information, you'll be asked to select the catalog you'd like to load into Select Star.

{% hint style="info" %}
💡 Select Star will not read queries or metadata or generate lineage for Catalogs, schemas, or tables that are not loaded. Please load all data for which you expect to see lineage.
{% endhint %}

You can [change the catalogs and schemas](https://docs.selectstar.com/data-source-management/manage-data-sources#configure-a-data-source) you have loaded if needed.

Select the catalogs and click **Next**.

<figure><img src="/files/RapsV8PCuymb1MAfXXAL" alt=""><figcaption></figcaption></figure>

For each catalog you selected, you'll be able to select the schemas.

<figure><img src="/files/EOScBnNVbuQhqqs2OmoU" alt=""><figcaption></figcaption></figure>

Your metadata should start loading automatically. Please allow 24-48 hours to completely generate popularity and lineage.

When the sync is complete, you'll be able to explore Databricks in Select Star.

See the link below for more information on Databricks in Select Star.

{% content-ref url="/pages/KGlgaHmYbTKOq1TVMuVR" %}
[Getting Started: Databricks](/learning-data/getting-started-databricks)
{% endcontent-ref %}


# BigQuery

Learn how to connect to your BigQuery data warehouse to Select Star and retrieve metadata from your datasets. This guide will provide you with the necessary steps to get started.

## Before you start

To connect BigQuery to Select Star, you will need...

* Admin access to your BigQuery instance.

Complete all of the following steps to see BigQuery metadata, lineage, and popularity in Select Star.

1. [Create a service account for Select Star](#id-1.-create-a-service-account-for-select-star)
2. [Create a new JSON key](#id-2.-create-a-new-json-key)
3. [Connect BigQuery to Select Star](#id-3.-connect-bigquery-to-select-star)
4. [Choose databases and schemas](#id-4.-choose-databases-and-schemas)

## 1. Create a Service Account for Select Star

Go to <https://console.cloud.google.com/iam-admin/serviceaccounts>.

Click **+ Create Service Account**.

![](/files/-MhL7QKJLXNqcFT4lYS4)

Grant the following roles to the service account:

* BigQuery Job User
* BigQuery Metadata Viewer
* BigQuery Resource Viewer
* BigQuery Data Viewer (optional: when this is turned off we won’t be able to show the detailed table information including table size, row count, etc)

![](/files/-MhL7QKKg7U1eNPh_1Vc)

Grant the <selectstar@getselectstar.com> user the Service account admin role.

![](/files/-MhL7QKLYaxLpCIHN_TF)

Click **Done**.

### Adding more projects

In order to add more projects in your Data Source you will need to have already created service account for one of the projects and:

1. Select your first project in which you have created the service account ![](/files/6oAOkHgDwWyv0MOak8bw)
2. Navigate to **IAM & Admin → Service accounts** in the project you have created the service account in and copy the email.

   <img src="/files/LMLJzmsUSZSF2Be2eEtl" alt="" data-size="original">

   **Copy the email** and save it for later.

   <img src="/files/6KFYECYzXm3sECNCkpsq" alt="" data-size="original">
3. Go to the **destination project**, i.e. the one that we want to grant the service account

   <img src="/files/quRMfl0Cjl4HASnUuhqo" alt="" data-size="original">
4. In **IAM & Admin → IAM** and click on **“ADD”** at the top.

   a. Select IAM in menu and follow the link.

   <img src="/files/ceCKIrgRMVaKNc8FMd5Z" alt="" data-size="original">

   b. Once in IAM, click "ADD"

   <img src="/files/Xf6eBMSoA7Ngk0Dcsu4u" alt="" data-size="original">
5. Use the email of service account from step 2 (that you copied) and grant the roles from [step 1](#2.-create-a-service-account-for-select-star)

   <img src="/files/nQKVeEIF4k8xBJ30iRwR" alt="" data-size="original">
6. You have successfully granted your service account permissions in another project, continue those steps if you want to add more projects.

## 2. Create a new JSON Key

Once the Service Account is created, click the Select Star Service Account in the list.

Go to **Keys** tab, and click **Add Key** → **Create new key.**

![](/files/-MhL7QKMrjevGvfr56K6)

Leave the **Key type** as `JSON`, and click **Create**. It will begin downloading the JSON file.

Save this JSON file to easily complete [step 3](#4.-connect-bigquery-to-select-star).

## 3. Connect BigQuery to Select Star

Go to the Select Star **Settings**. Click **Data** in the sidebar, then **+ Add** to create a new Data Source.

![](/files/IP5pzA0NjiBfKRqyS8wu)

Choose **BigQuery** in the Source Type dropdown and provide the following information:

![](/files/u6fHutnsu2cHqjw6DZ4i)

* **Display Name:** This value is `BigQuery` by default, but you can override it if desired.
* **Service Account:** The JSON file created in [step 2](#3.-create-a-new-json-key) above. The JSON file will automatically populate all other fields in the modal.

## 4. Choose databases and schemas

After you fill in the information, you'll be asked to select the databases you'd like to load into Select Star.

{% hint style="info" %}
Select Star will not read queries or metadata or generate lineage for databases, schemas, or tables that are not loaded. Please load all data for which you expect to see lineage.

You can [change the databases and schemas](/data-source-management/manage-data-sources#configure-a-data-source) you have loaded if needed.
{% endhint %}

Select the database and click **Next**.

![](/files/49WfnJ4yST1arpDJXuYw)

For each database you selected, you'll be able to select the schemas.

![](/files/UIi8Sawb1QHQ6mMS99tQ)

Your metadata should start loading automatically. Please allow 24-48 hours to completely generate popularity and lineage.

When the sync is complete, you'll be able to explore your BigQuery instance in Select Star. See the link below for more information on BigQuery in Select Star.

{% content-ref url="/pages/-MgSY3JsJz5uuNb23QXn" %}
[Getting Started: BigQuery](/learning-data/getting-started-bigquery)
{% endcontent-ref %}

{% content-ref url="/pages/7YL53MsYRajOkgrlKRGc" %}
[BigQuery Description Sync (beta)](/integrations/bigquery/bigquery-description-sync)
{% endcontent-ref %}


# BigQuery Description Sync (beta)

You have the option of syncing descriptions created in Select Star back to BigQuery. This beta feature allows you to push back descriptions for schemas, tables, and columns to your original BigQuery data source.

{% hint style="info" %}
This feature is currently in beta and only available for BigQuery data sources.
{% endhint %}

## Enable BigQuery Description Sync in Select Star

To try out this feature while it is in beta, contact Select Star support.

{% hint style="info" %}
The service account you created for Select Star must have the necessary permissions to update BigQuery metadata. Ensure your service account has the `bigquery.tables.update` and `bigquery.datasets.update` permissions.

To grant these permissions, assign the appropriate IAM roles to your service account in the Google Cloud Console. You can do this by navigating to **IAM & Admin > IAM**, finding your service account, and editing its permissions. For more details, see the [Google Cloud documentation on granting roles and permissions](https://cloud.google.com/iam/docs/granting-changing-revoking-access).
{% endhint %}

## How Description Sync Works

When Description Sync is enabled:

* Descriptions added to schemas, tables, and columns in Select Star will be synced back to BigQuery. This sync is performed automatically, once every day.
* The Descriptions added by users are synced back. Descriptions from other sources (e.g. AI generated, or propagated from other objects) are only synced back if they are explicitly selected as the description for the object.
* The sync process updates the metadata descriptions in your BigQuery data source and the changes are reflected in BigQuery's information schema and console interface.
* Changes to descriptions in BigQuery will still be reflected in Select Star.


# AWS Redshift

Follow these steps to connect your AWS Redshift instance to Select Star.

## Before you start

To connect AWS Redshift to Select Star, you will need...

* access to CloudFormation with permissions to modify IAM, AWS Lambda, Redshift cluster and VPC
* access to AWS Redshift admin

Our Amazon Redshift integration is currently designed only for the provisioned (cluster-based) deployment type and uses the standard Redshift API.

{% hint style="info" %}
Select Star requires only minimal metadata access to AWS Redshift. The granted permissions are defined in CloudFormation template:

* IAM permission defined by resource "CrossAccountRolePolicy" in file [SelectStarRedshift.json](https://github.com/selectstar/cloudformation-templates/blob/main/redshift/SelectStarRedshift.json)
* AWS Redshift user permission defined in file [provision.py](https://github.com/selectstar/cloudformation-templates/blob/main/redshift/provision.py)

For instances operating within a non-publicly accessible environment, such as an AWS VPC, please refer to our guide on [Integrating Private Network Data Sources](/integrations/private-network) for detailed instructions and best practices.
{% endhint %}

1. [Create a new data source](#1.-create-a-new-data-source)
2. [Create CloudFormation stack](#2.-create-cloudformation-stack)
3. [Confirm authorization](#3.-confirm-authorization)
4. [Choose databases and schemas](#4.-choose-databases-and-schemas)

## 1. Create a new data source

1\. Go to the Select Star **Settings**. Click **Data** in the sidebar, then **+ Add** to create a new Data Source.

![Animations shows where to click to create a new data source.](/files/IP5pzA0NjiBfKRqyS8wu)

2\. Fill form in the required information:

* **Source Type:** Select "Redshift"
* **Display Name:** This value is `Redshift` by default, but you can override it if desired.
* **Cluster Name:** The name of your AWS Redshift cluster in the AWS management console. Also known as "Cluster identifier" by AWS.
* **Database:** The name of the database in AWS Redshift you've given us access to.
* **AWS Region:** ID of the AWS region where the cluster was created. For example `us-east-2`,`us-west-1`, `eu-central-1`

![Screenshot shows a new data source form.](/files/6jHUujSY4GIU3wS7ylSZ)

3\. Click **Connect**.

## 2. Create CloudFormation stack

Select Star recommends use of AWS CloudFormation to setup integration, which allows you to make necessary changes to the Redshift cluster environment in a automatic, transparent, safe and auditable manner.

AWS CloudFormation template create AWS resources, modify and validate the Redshift configuration for safe integration with the Select Star services:

* validate Redshift cluster compatibility
* create an AWS IAM Role to enable access for Select Star and add it to Redshift cluster
* create S3 bucket for query logs
* activates log export to S3 bucket of all Redshift logs, if log export is not enabled so far
* create a custom parameter group or modify an existing one to set a parameter `enable_user_activity_logging`
* create Redshift user `selectstar` and grants permissions in Redshift cluster
* configure the security group to allow Select Star access to the Redshift cluster

By default, AWS CloudFormation template activates log export to AWS S3. To use AWS CW, [activate exporting to AWS CW beforehead](https://docs.aws.amazon.com/redshift/latest/mgmt/db-auditing.html).

The source code of the CloudFormation template along with build scripts and real-time logs of the continuous deployment system is available on public repository on GitHub "[selectstar/cloudformation-templates](https://github.com/selectstar/cloudformation-templates/tree/main/redshift)" to be freely audited.

You can pass the link to CloudFormation to the infrastructure team to enable the integration to be created.

1\. A simple form will be displayed in Select Star.

![Screenshot shows the access authorization form.](/files/LHLk95d3mFJrt5rumbJo)

2\. Select the "Open CloudFormation" button. A new window will open to proceed to the creation of a CloudFormation stack by AWS Management Console. Make sure you are logged into the AWS account in which the Redshift cluster exist.

3\. The **Create Stack** form will be displayed. Fill form in the required information:

* Under Parameters, enter the Amazon Redshift cluster name, list of comma separated database names, and your database user name. The indicated user will be used only by CloudFormation to create a dedicated user "select\_star". Select Star will not have any access to the indicated user.
* Select "true" in the "Configure S3 logging" and "Restart Cluster (if necessary to apply changes)" fields for fully automatic cluster configuration.

{% hint style="info" %}
Select Star will **only have access to the databases that exist at the time of provisioning**. If you are planning to add new databases at a later stage, we suggest you do that before provisioning all the permissions.
{% endhint %}

{% hint style="warning" %}
The user you provide (**DbUser**) **needs to have admin access**. This is the user through which the CloudFormation template will create the user that Select Star needs with minimal access.
{% endhint %}

<figure><img src="/files/SvuOPEThxXEMexIqySuy" alt="" width="375"><figcaption><p>Redshift configuration</p></figcaption></figure>

4\. Review the information and under **Capabilities** choose "I acknowledge that AWS CloudFormation might create IAM resources".

<div align="center"><img src="/files/rG3RZbmah9RiXCYbOvGq" alt="Screenshot shows the &#x22;Capabilities&#x22; section in &#x22;Quick create stack&#x22; form." width="375"></div>

5\. Choose **Create stack**.

6\. Wait until the stack changes it status to "CREATE\_COMPLETE" from "CREATE\_IN\_PROGRESS" in tab "Stack info". The operation should take up to 5 minutes. You need to refresh tab to see the progress.

![Screenshot shows the "State" value in "Stack info" tab.](/files/nGoWjGcsGrvasNJi9VZS)

7\. After completing stack creation, the `Role ARN` is available from the "Outputs". Copy and save the `RoleArn` for later use.

![Screenshot shows where to obtain an Role ARN in the Outputs tab](/files/jXbcSSwdPBTkEhUZkGkX)

## 3. Confirm authorization

1\. Return to Select Star. You should see a form that allows you to provide "Role ARN". Fill form in the required information:

* **Role ARN:** Identifier of AWS IAM Role to use by Select Star. You'll see this after completing [step 2.7](#2.-create-cloudformation-stack) of the instructions.

![Screenshot shows authroization in Select Star](/files/KkwMOekodbwI0tZfgFDM)

2\. Click **Connect**.

## 4. Choose databases and schemas

After you fill in the information, you'll be asked to select the databases you'd like to load into Select Star.

{% hint style="info" %}
Select Star will not read queries or metadata or generate lineage for databases, schemas, or tables that are not loaded. Please load all data for which you expect to see lineage.

You can [change the databases and schemas](https://docs.selectstar.com/data-source-management/manage-data-sources#configure-a-data-source) you have loaded if needed.
{% endhint %}

Select the database and click **Next**.

![Database selection](/files/-MkdC0euUkCjrubKDgis)

For each database you selected, you'll be able to select the schemas.

![Schema selection](/files/-MkdC0evROqyuF58x_bx)

Your metadata should start loading automatically. Please allow 24-48 hours to completely generate popularity and lineage.

When the sync is complete, you'll be able to explore Redshift in Select Star.


# Manual setup

This guide provides comprehensive instructions for manually setting up Redshift integration with Select Star. The setup process includes creating an IAM role, configuring Redshift logging, managing databases and users, and establishing network access. Follow these steps carefully to ensure a successful integration.

Use this manual process if automatic integration using the AWS CloudFormation template is not possible. However, it is strongly recommended to [use the AWS CloudFormation template setup](/integrations/redshift) to minimize manual errors and ensure all necessary steps are completed efficiently.

1. [Create a new data source](#id-1.-create-a-new-data-source)
2. [Access AWS Console](#id-2.-access-aws-console)
3. [Check Cluster State](#id-3.-check-cluster-state)
4. [Logging Configuration](#id-4.-logging-configuration)
5. [Cluster Parameter Group](#id-5.-cluster-parameter-group)
6. [Enable User Activity Logging](#id-6.-enable-user-activity-logging)
7. [Reboot Cluster](#id-7.-reboot-cluster)
8. [Database and User Management](#id-8.-database-and-user-management)
9. [IAM Role Management](#id-9.-iam-role-management)
10. [Network Management](#id-10.-network-management)
11. [Confirm authorization](#id-11.-confirm-authorization)
12. [Choose databases and schemas](#id-12.-choose-databases-and-schemas)

## 1. Create a new data source

1\. Go to the Select Star **Settings**. Click **Data** in the sidebar, then **+ Add** to create a new Data Source.

![Animation shows where to click to create a new data source.](/files/IP5pzA0NjiBfKRqyS8wu)

2\. Fill in the required information:

* *z*Source Type:\*\* Select "Redshift"
* **Display Name:** This value is `Redshift` by default, but you can override it if desired.
* **Cluster Name:** The name of your AWS Redshift cluster in the AWS management console. Also known as "Cluster identifier" by AWS.
* **Database:** The name of the database in AWS Redshift you've given us access to.
* **AWS Region:** ID of the AWS region where the cluster was created. For example `us-east-2`, `us-west-1`, `eu-central-1`.

![Screenshot shows a new data source form.](/files/6jHUujSY4GIU3wS7ylSZ)

3\. Click **Connect**.

## 2. Access AWS Console

1. **Log into the AWS Management Console**
   * Ensure you have the necessary permissions to manage Redshift clusters and VPC.

## 3. Check Cluster State

1. **Ensure the Cluster is Available**
   * The cluster should be publicly accessible. Modify the cluster if necessary.
   * For instances operating within a non-publicly accessible environment, such as an AWS VPC, please refer to our guide on [Integrating Private Network Data Sources](/integrations/private-network) for detailed instructions and best practices.

## 4. Logging Configuration

1. **In the “Properties” Tab Scroll to the “Database Configurations” Section**
2. **Enable Audit Logging**:
   * Use the "Edit" button to enable audit logging by specifying an S3 bucket for logs.

For details, refers to AWS documentation about [database audit logging](https://docs.aws.amazon.com/redshift/latest/mgmt/db-auditing.html).

## 5. Cluster Parameter Group

1. **If the Current Parameter Group is Default, Create a New Custom Parameter Group**
   * Base it on the family of the existing parameter group.
2. **Apply This New Parameter Group to Your Cluster**

For details, refers to AWS documentation about [managing parameter groups using the console](https://docs.aws.amazon.com/redshift/latest/mgmt/managing-parameter-groups-console.html).

## 6. Enable User Activity Logging

1. **Modify the Custom Parameter Group**
   * Set the `enable_user_activity_logging` parameter to `true`.

## 7. Reboot Cluster

1. **Reboot the Cluster if Changes Require It**
   * Operations such as switching the parameter group or modifying its parameters may require a restart to apply.
   * If required by your organization's management policy, ensure to schedule downtime before restarting the cluster.

## 8. Database and User Management

1. **Go to the “Query Editor” in Redshift**
2. **Run SQL Commands to Create Users or Grant Permissions**:

   ```sql
   CREATE USER selectstar WITH PASSWORD DISABLE syslog ACCESS UNRESTRICTED;
   GRANT SELECT ON SVV_TABLE_INFO TO selectstar;
   GRANT SELECT ON SVV_TABLES TO selectstar;
   GRANT SELECT ON SVV_COLUMNS TO selectstar;
   GRANT SELECT ON STL_QUERYTEXT TO selectstar;
   GRANT SELECT ON STL_DDLTEXT TO selectstar;
   GRANT SELECT ON STL_QUERY TO selectstar;
   ```

   Retry operation in context of each database which should be ingested in SelectStar

## 9. IAM Role Management

1\. **Navigate to the AWS IAM Console** 2. **Create a new IAM Role** 3. **Note ARN of AWS IAM Role**

* It will be required in Select Star UI. 4. **Use In-line Permission Policy**:

```json
{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Action": [
                "redshift:DescribeClusters",
                "redshift:DescribeLoggingStatus",
                "redshift:ListSchemas",
                "redshift:ListTables",
                "redshift:ListDatabases",
                "redshift:ExecuteQuery",
                "redshift:FetchResults",
                "redshift:CancelQuery",
                "redshift:DescribeQuery",
                "redshift:DescribeTable",
                "redshift:ViewQueriesFromConsole"
            ],
            "Resource": "arn:aws:redshift:{region_id}:{account_id}:cluster:{cluster_name}",
            "Effect": "Allow"
        },
        {
            "Action": [
                "redshift:GetClusterCredentials"
            ],
            "Resource": "arn:aws:redshift:{region_id}:{account_id}:dbuser:{cluster_name}/selectstar",
            "Effect": "Allow"
        },
        {
            "Action": [
                "redshift:GetClusterCredentials"
            ],
            "Resource": [
                "arn:aws:redshift:{region_id}:{account_id}:dbname:{cluster_name}/*"
            ],
            "Effect": "Allow"
        },
        {
            "Action": [
                "redshift-data:ExecuteStatement",
                "redshift-data:ListDatabases",
                "redshift-data:ListSchemas",
                "redshift-data:ListTables",
                "redshift-data:DescribeTable"
            ],
            "Resource": "arn:aws:redshift:{region_id}:{account_id}:cluster:{cluster_name}",
            "Effect": "Allow",
            "Sid": "DataAPIPermissions"
        },
        {
            "Action": [
                "redshift-data:GetStatementResult",
                "redshift-data:CancelStatement",
                "redshift-data:DescribeStatement",
                "redshift-data:ListStatements"
            ],
            "Resource": "*",
            "Effect": "Allow",
            "Sid": "DataAPIIAMSessionPermissionsRestriction"
        },
        {
            "Action": [
                "s3:GetLifecycleConfiguration",
                "s3:GetBucketTagging",
                "s3:GetInventoryConfiguration",
                "s3:GetObjectVersionTagging",
                "s3:ListBucketVersions",
                "s3:GetBucketLogging",
                "s3:GetBucketPolicy",
                "s3:GetBucketOwnershipControls",
                "s3:GetBucketPublicAccessBlock",
                "s3:GetBucketPolicyStatus",
                "s3:ListBucketMultipartUploads",
                "s3:GetBucketVersioning",
                "s3:GetBucketAcl",
                "s3:ListMultipartUploadParts",
                "s3:GetObject",
                "s3:GetBucketLocation",
                "s3:GetObjectVersion",
                "s3:ListBucket"
            ],
            "Resource": [
                "arn:aws:s3:::{bucket}",
                "arn:aws:s3:::{bucket}/*"
            ],
            "Effect": "Allow",
            "Sid": "ListObjectsInBucket"
        }
    ]
}
```

Replace placeholders like `{cluster_name}`, `{region_id}`, `{account_id}`, `{bucket}`, etc., with your specific details.

5\. **Use for AWS IAM Role Assume Policy**:

```json
{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Effect": "Allow",
            "Principal": {
                "AWS": "{select_star_principal_id}"
            },
            "Action": "sts:AssumeRole",
            "Condition": {
                "StringEquals": {
                    "sts:ExternalId": "{external_id}"
                }
            }
        }
    ]
}
```

Replace placeholders like `{select_star_principal_id}` and `{external_id}` with values specific for your data source instance. Please contact Select Star support to obtain values after creating the data source. They are also encoded in the link "Open CloudFormation" in Select Star UI.

## 10. Network Management

To create an AWS VPC ingress rule for your Redshift cluster to allow access from specific IP addresses of the Select Star platform (3.23.108.85/32 and 3.20.56.105/32) via the AWS Management Console, follow these steps:

1. **Navigate to the VPC Console**: In the AWS Management Console, go to "Services" and select "VPC".
2. **Open Security Groups**: In the VPC dashboard, locate and click on "Security Groups" in the left-hand navigation pane.
3. **Find the Redshift Cluster’s Security Group**: If you know the Security Group ID associated with your Redshift cluster, find it directly. Otherwise, navigate to the Redshift dashboard, select your cluster, and note the Security Group under its properties.
4. **Select the Security Group**: Click on the relevant Security Group for your Redshift cluster.
5. **Switch to the “Inbound Rules” Tab**: Click on the “Inbound rules” tab.
6. **Click on “Edit Inbound Rules”**: Click the "Edit inbound rules" button.
7. **Add a Rule for the First IP**:
   * Click “Add Rule”.
   * For “Type”, select “Redshift” (default port is TCP 5439).
   * Select "Custom" in "Source" and enter `3.23.108.85/32`.
   * Optionally, add a description.
8. **Add a Rule for the Second IP**:
   * Click “Add Rule” again.
   * Repeat the same steps but enter `3.20.56.105/32` in “Source”.
9. **Review and Save the Rules**: Ensure the rules are correctly set up, then click “Save rules” to apply the new inbound rules.

Always confirm that you are modifying the correct security group associated with your Redshift cluster to avoid any unintended access issues.

## 11. Confirm Authorization

1\. Return to Select Star. You should see a form that allows you to provide the "Role ARN". Fill in the required information:

* **Role ARN:** Identifier of the AWS IAM Role to use by Select Star. You'll see this after creating the new IAM Role.

![Screenshot shows authroization in Select Star](/files/KkwMOekodbwI0tZfgFDM)

2\. Click **Connect**.

## 12. Choose Databases and Schemas

After you fill in the information, you'll be asked to select the databases you'd like to load into Select Star.

{% hint style="info" %}
Select Star will not read queries or metadata or generate lineage for databases, schemas, or tables that are not loaded. Please load all data for which you expect to see lineage.

You can [change the databases and schemas](https://docs.selectstar.com/data-source-management/manage-data-sources#configure-a-data-source) you have loaded if needed.
{% endhint %}

Select the database and click **Next**.

![Database selection](/files/-MkdC0euUkCjrubKDgis)

For each database you selected, you'll be able to select the schemas.

![Schema selection](/files/-MkdC0evROqyuF58x_bx)

Your metadata should start loading automatically. Please allow 24-48 hours to completely generate popularity and lineage.

When the sync is complete, you'll be able to explore Redshift in Select Star.

## Notes

* Replace placeholders like `{cluster}`, `{role}`, `{bucket}`, `{db}`, etc., with your specific details.
* Some steps may require waiting for the cluster to become available after changes.
* Always verify changes and ensure they align with your infrastructure and security policies.


# Microsoft SQL Server / MS SQL (beta)

Follow these steps to connect your Microsoft SQL Server instance to Select Star.

1. [Create a Service Account](#id-1.-create-a-service-account)
2. [Grant additional permissions for preview access. (optional)](#id-2.-grant-additional-permissions-for-preview-access.-optional)
3. [Create a new datasource in Select Star](#id-3.-create-a-new-data-source-in-select-star)
4. [Enable query logging](#id-4.-enable-query-logging)
5. [Choose databases and schemas](#id-5.-choose-databases-and-schemas)

For instances operating within a non-publicly accessible environment, such as a private network, please refer to our guide on [Integrating Private Network Data Sources](/integrations/private-network) for detailed instructions and best practices.

## 1. Create a Service Account

We recommend creating a dedicated service account for Select Star with metadata-only roles and permissions.

Execute the following commands:

```sql
CREATE LOGIN <username> WITH PASSWORD = '<password>';
CREATE USER <username> FOR LOGIN <username>;
GRANT VIEW DEFINITION TO <username>;
```

### For Microsoft SQL Server

Execute the following additional commands:

```sql
-- Run use master and GRANT together,
use master
GRANT VIEW SERVER STATE TO <username>;
```

### For Cloud SQL for SQL Server

Log in as `sqlserver` user and execute the following command:

```sql
GRANT VIEW SERVER STATE TO <username> AS CustomerDbRootRole;
```

## 2. Grant additional permissions for preview access. (optional)

{% hint style="info" %}
This step is required if you wish to enable Select Star's **Preview** Feature.
{% endhint %}

Execute the following command:

```sql
-- Grant access to all tables
GRANT SELECT TO <username>;
```

## 3. Create a new data source in Select Star

* **Display Name** - The Name you want to give to your new datasource.
* **Hostname** - Your hostname defines the location where your database is hosted.
* **Port** - The communication endpoint used to connect clients to the SQL Server instance.
* **Username** - Specify the User to connect to the SQL Server instance. It should have enough privileges to read all the metadata.
* **Password** - Password.
* **Database** - The database name.

![Create Data Source](/files/eJTuJmBPiJCRGxTyxsuI)

{% hint style="info" %}
Select Star currently does not support Popularity and Lineage is only supported for views in this integration. If needed, please feel free to contact us so we can explore alternative solutions.
{% endhint %}

## 4. Enable query logging

Select Star platform supports various query log sources, for optimal selection, please refer to [the documentation](/integrations/mssql/query-log).

## 5. Choose databases and schemas

After you fill in the information, you'll be asked to select the databases you'd like to load into Select Star.

{% hint style="info" %}
Select Star will not read queries or metadata for databases, schemas, or tables that are not loaded. Please load all data for which you expect to see metadata.

You can [change the databases and schemas](/data-source-management/manage-data-sources#configure-a-data-source) you have loaded if needed.
{% endhint %}

Your metadata should start loading automatically, once the sync is complete, you'll be able to see it in Select Star.

See the link below for more information on how to navigate through data warehouses.

{% content-ref url="/pages/Upyz3faFJpQ1KvNypm1x" %}
[Getting Started: Data Warehouse](/learning-data/getting-started-data-warehouse)
{% endcontent-ref %}


# Query Logs

Available query logs sources for Microsoft SQL data sources:

* System tables / views
* Database Activity Streams on an Amazon RDS instance

{% hint style="info" %}
If you are using a different hosting provider and you want to ingest query logs, please contact our sales team.
{% endhint %}

## Database Activity Streams on an Amazon RDS instance

To enable Database Activity Streams on an Amazon RDS instance running Microsoft SQL Server (MSSQL), follow the step-by-step instructions below using the AWS Management Console.

### Step 1: Create a Database Audit Specification

Execute query on database instance to create a database audit specification called `RDS_DAS_DB_SELECT_STAR_QUERY_STREAM` and associate it with the server audit `RDS_DAS_AUDIT`, ensuring that all matched events are sent to the Kinesis DAS. The specified policy will capture any of the following command types: `SELECT`, `UPDATE`, `INSERT`, `DELETE`, `EXECUTE`, `RECEIVE`, and `REFERENCES` on the `olist` database, but only for commands submitted by the public principal (public matches all users). The policy will be activated immediately with `WITH (STATE = ON)`.

The query:

```sql
CREATE DATABASE AUDIT SPECIFICATION [RDS_DAS_DB_SELECT_STAR_QUERY_STREAM]
FOR SERVER AUDIT [RDS_DAS_AUDIT]
ADD (
  SELECT, UPDATE, INSERT, DELETE, EXECUTE, RECEIVE, REFERENCES
  ON Database::olist
  BY public
)
WITH (STATE = ON);
```

To enable query log for multiple databases, execute operation in each database individually. In SQL Server, including the version running on AWS RDS, there is no direct way to log queries across all databases without setting up database-level audit specifications for each database individually. SQL Server auditing is designed to be granular, and the audit specifications must be created at the database level to capture events occurring within that specific database.

### Step 2: Sign in to the AWS Management Console

1. Open the AWS Management Console at <https://console.aws.amazon.com/rds/>.
2. Sign in with your AWS credentials.

### Step 3: Navigate to the RDS Dashboard

1. In the AWS Management Console, search for **RDS** in the top search bar.
2. Click on **RDS** to open the Amazon RDS Dashboard.

### Step 4: Select Your RDS Instance

1. In the RDS Dashboard, click on **Databases** in the left-hand menu.
2. Locate your Microsoft SQL Server RDS instance in the list of databases.
3. Click on the **DB Identifier** of the MSSQL instance you wish to enable the Database Activity Stream for.
4. Verify DB instance class used by instance with [supported DB instance classes for database activity streams](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/DBActivityStreams.html#DBActivityStreams.Overview.requirements.classes)

### Step 5: Start the Database Activity Stream

1. On the database details page, click on the **Actions** dropdown menu located in the top-right corner.
2. Select **Start activity stream** from the dropdown menu.

### Step 6: Configure the Database Activity Stream

1. The **Start database activity stream: name** window will appear, where `name` is your RDS instance.
2. Configure the following settings:
   * **AWS KMS key**: Choose a KMS key from the list of available AWS KMS keys. This key is used to encrypt the key that encrypts the database activity stream data.
   * **Database activity events**: Check the box to **Enable engine-native audit fields**.
3. Choose **Immediately** if you can restart the RDS instance right away and start the database activity stream immediately. If you select **During the next maintenance window**, the RDS instance will restart during the next scheduled maintenance window, and the activity stream will start at that time.

You can refer to AWS documentation about [starting a database activity stream](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/DBActivityStreams.Enabling.html) for details.

Entries in the activity stream are encrypted using unique batch keys before being written to Kineses. To facilitate decryption, the session keys is itself encrypted using the selected KMS key and included in the batch. Anyone with access to the KMS key will be able to decrypt the batch keys and use those to access raw audit events from the database.

### Step 7: Start the Activity Stream

1. After configuring the settings, click on **Start database activity stream**.
2. The status of the database will indicate that the activity stream is starting.

### Step 8: Locate AWS KMS key and AWS Kinesis stream

1. Once on the database details page, open the **Configuration** tab.
2. In the **Configuration** tab, locate in section **Database activity stream** following:
   * **AWS KMS key**: This shows the current KMS key used by the Database Activity Stream. Note this key may be different one that one used for the RDS instance.
   * **Kinesis data stream**: Note the configured Kinesis stream.

### Step 9. Create CloudFormation stack

Select Star recommends setting up integration using AWS CloudFormation, which allows you to make necessary changes to the RDS instance environment automatically, transparently, safely, and auditably. It deploys AWS Firehose and AWS Lambda function to store relevant information in AWS S3 bucket and enable us to access the AWS S3 bucket.

The source code of the CloudFormation template, build scripts, and real-time logs of the continuous deployment system are available on our public repository [selectstar/cloudformation-templates](https://github.com/selectstar/cloudformation-templates/tree/main/das), hosted on Github, to be freely audited.

You can pass the link to CloudFormation to the infrastructure team to enable the integration to be created.

1\. A simple form will be displayed in Select Star. In option **Source Type** select "RDS Database Activity Stream".

2\. Click **Open CloudFormation**. A new window will open to proceed to the creation of a CloudFormation stack by AWS Management Console. Make sure you are logged into the AWS account in which the RDS instance is hosted.

3\. The **Create Stack** form will be displayed. Some of the values will be filled in by default, under *Parameters*, enter:

* **KinesisStreamARN**: The ARN of the Kinesis stream where RDS Database Activity Streams are delivered. Example: `arn:aws:kinesis:us-east-2:000000000:stream/aws-rds-das-db-ZQO7M43PGGUXJEZVYSALTO76KA`
* **KmsKeyARN**: The ARN of the KMS key used for encryption of RDS Database Activity Streams. Example: `arn:aws:kms:us-east-2:000000000:key/f319545f-a0d4-4bfc-896f-5d37fe921ffb`
* **RdsResourceId**: The ARN of the RDS instance or cluster producing the logs. Example: `db-ZQO7M43PGGUXJEZVYSALTO76KA`

4\. Review the information and under **Capabilities** choose "I acknowledge that AWS CloudFormation might create IAM resources".

![Screenshot shows the "Capabilities" section in "Quick create stack" form.](/files/jCCAaTjHQ6eD4FWu984f)

5\. Choose **Create stack**.

6. Wait until the stack changes its status to "CREATE\_COMPLETE" from "CREATE\_IN\_PROGRESS " in the tab "Stack info." The operation should take up to 5 minutes. You need to refresh the tab to see the progress.

![Screenshot shows the "State" value in "Stack info" tab.](/files/ElcWSxmde3BEveYyvznJ)

7\. After completing stack creation, the `Role ARN` and `S3BucketName` is available from the "Outputs". Copy and save the `RoleArn` and `S3BucketName` for later use.

8\. Logs are published in batches, which can be buffered up to 15 minutes. After configuring and executing a database query, new files should appear in the bucket under the `processed` prefix.

### Step 10. Confirm authorization

1\. Return to Select Star. You should see a form that allows you to provide "Role ARN" . Fill form in the required information:

* **Type**: Select "RDS Database Activity Stream"
* **AWS region**: Identifier of the AWS Region where your MSSQL instance is hosted eg. `us-east-2`
* **Role ARN:** Identifier of the AWS IAM Role to use by Select Star. Available in the "Outputs" tab of AWS CloudFormation.
* **S3 bucket:** The name of the AWS S3 bucket which stores query logs. Available in the "Outputs" tab of AWS CloudFormation.
* **S3 object prefix:** The name of the AWS S3 prefix which stores query logs. For deployment using AWS CloudFormation, this is `processed/` by default.

2\. Click **Connect**.


# MySQL (beta)

Follow these steps to connect your MySQL instance to Select Star.

1. [Create a Service Account](#id-1.-create-a-service-account)
2. [Create a new datasource in Select Star](#id-2.-create-a-new-data-source-in-select-star)
3. [Enable query logging](#id-3.-enable-query-logging)
4. [Choose databases and schemas](#id-4.-choose-databases-and-schemas)

For instances operating within a non-publicly accessible environment, such as a private network, please refer to our guide on [Integrating Private Network Data Sources](/integrations/private-network) for detailed instructions and best practices.

## 1. Create a Service Account

We recommend creating a dedicated service account for Select Star with read-only access to metadata and data.

{% hint style="info" %}
MySQL requires `SELECT` to execute `DESCRIBE`, `SHOW CREATE TABLE`, and queries on `INFORMATION_SCHEMA`, which retrieve table metadata. Without it, access to table structure is denied. `SHOW VIEW` is needed for views.
{% endhint %}

Execute the following command:

```sql
CREATE USER '<username>' IDENTIFIED BY '<password>';
GRANT SELECT ON *.* TO '<username>'
GRANT SHOW VIEW ON *.* TO '<username>'

-- Apply the changes
FLUSH PRIVILEGES;
```

## 2. Create a new data source in Select Star

* **Display Name** - The Name you want to give to your new datasource.
* **Hostname** - Your hostname defines the location where your database is hosted.
* **Port** - The communication endpoint used to connect clients to the MySQL instance.
* **Username** - Specify the User to connect to. It should have enough privileges to read all the metadata.
* **Password** - Password to connect the MySQL instance.

![Create Data Source](/files/p62yVc7eUeNGlAwUsAnP)

{% hint style="info" %}
Select Star currently does not support Popularity and Lineage for this integration. If needed, please feel free to contact us so we can explore alternative solutions.
{% endhint %}

## 3. Enable query logging

Select Star platform supports various query log sources, for optimal selection, please refer to [the documentation](/integrations/mysql/query-log).

### 4. Choose databases and schemas

After you fill in the information, you'll be asked to select the databases you'd like to load into Select Star.

{% hint style="info" %}
Select Star will not read queries or metadata for databases, schemas, or tables that are not loaded. Please load all data for which you expect to see metadata.

You can [change the databases and schemas](/data-source-management/manage-data-sources#configure-a-data-source) you have loaded if needed.
{% endhint %}

Your metadata should start loading automatically, once the sync is complete, you'll be able to see it in Select Star.

See the link below for more information on how to navigate through data warehouses.

{% content-ref url="/pages/Upyz3faFJpQ1KvNypm1x" %}
[Getting Started: Data Warehouse](/learning-data/getting-started-data-warehouse)
{% endcontent-ref %}


# Query Logs

Available query logs sources for MySQL data sources:

* None
* CloudWatch Logs

{% hint style="info" %}
If you are using a different hosting provider and you want to ingest query logs, please contact our sales team.
{% endhint %}

## CloudWatch Logs

For instances hosted on AWS RDS we have the ability to ingest query log which allows us to get detailed lineage and popularity. Here's a step-by-step guide to enable the General Query Log for a MySQL database hosted on AWS RDS and export the logs to AWS CloudWatch.

### Step 1: Sign in to the AWS Management Console

1. Go to the [AWS Management Console](https://aws.amazon.com/console/).
2. Sign in using your AWS account credentials.

### Step 2: Navigate to the RDS Dashboard

1. In the AWS Management Console, search for "RDS" in the search bar at the top and select **RDS** from the search results.
2. This will take you to the RDS dashboard.

### Step 3: Select Your RDS Instance

1. In the RDS dashboard, click on **Databases** in the left-hand menu.
2. Find and select the MySQL RDS instance for which you want to enable the General Query Log.

### Step 4: Modify the DB Parameter Group

1. On your selected DB instance page, scroll down to find the **Configuration** section.
2. Under **Parameter group**, note the name of the currently associated parameter group. If it's a default parameter group, you need to create a new custom parameter group because default ones cannot be modified.
3. In the left-hand menu, click on **Parameter groups**.
4. Click **Create parameter group** if you need to create a new one:
   * Choose the **Parameter group family** corresponding to your MySQL version.
   * Enter a name for the new parameter group.
   * Click **Create**.
5. Once created, find your new parameter group in the list, and click on it to modify the parameters.

### Step 5: Enable General Query Log in the Parameter Group

1. In the parameter group settings, use the search bar to find the `general_log` parameter.
2. Set `general_log` to `1` (enabled).
3. Find the `log_output` parameter and set it to `FILE` as required for logging to CloudWatch.
4. Save the changes.

### Step 6: Apply the Parameter Group to Your RDS Instance

1. Go back to the **Databases** section.
2. Select your MySQL RDS instance again.
3. Click on the **Modify** button in the upper right corner.
4. In the **Database options** section, change the **DB Parameter Group** to the new custom parameter group you just modified.
5. Scroll down and choose whether to apply the changes immediately or during the next maintenance window.
6. Click **Continue** and then **Modify DB Instance**.

### Step 7: Enable Logging to CloudWatch

1. With the RDS instance selected, go to the **Logs & events** tab.
2. Under **Manage export logs to CloudWatch**, ensure that **General Log** is checked.
3. Click on **Configure**.

### Step 8: Confirm Logs in CloudWatch

1. Go to the AWS Management Console and search for **CloudWatch**.
2. In the CloudWatch dashboard, click on **Logs** in the left-hand menu.
3. You should see a log group corresponding to your RDS instance. The General Query Log entries should now appear here.

### Step 9: Test and Verify

1. Run a few queries on your MySQL database.
2. Check CloudWatch Logs to ensure that the queries are being logged and that the General Query Log is working correctly.

### Step 10. Create CloudFormation stack

Select Star recommends setting up integration using AWS CloudFormation, which allows you to make necessary changes to the RDS instance environment automatically, transparently, safely, and auditably.

AWS CloudFormation will create AWS resources to enable us to access AWS CloudWatch logs.

The source code of the CloudFormation template, build scripts, and real-time logs of the continuous deployment system are available on our public repository [selectstar/cloudformation-templates](https://github.com/selectstar/cloudformation-templates/tree/main/mysql), hosted on Github, to be freely audited.

You can pass the link to CloudFormation to the infrastructure team to enable the integration to be created.

1\. A simple form will be displayed in Select Star. In option **Source Type** select "CloudWatch Logs".

2\. Click **Open CloudFormation**. A new window will open to proceed to the creation of a CloudFormation stack by AWS Management Console. Make sure you are logged into the AWS account in which the RDS instance is hosted.

3\. The **Create Stack** form will be displayed. Some of the values will be filled in by default, under *Parameters*, enter:

* **Log Group Name**: The name of the log group to be consumed by Select Star. Example: `/aws/rds/instance/dev-mysql8/general`

4\. Review the information and under **Capabilities** choose "I acknowledge that AWS CloudFormation might create IAM resources".

![Screenshot shows the "Capabilities" section in "Quick create stack" form.](/files/jCCAaTjHQ6eD4FWu984f)

5\. Choose **Create stack**.

6. Wait until the stack changes its status to "CREATE\_COMPLETE" from "CREATE\_IN\_PROGRESS " in the tab "Stack info." The operation should take up to 5 minutes. You need to refresh the tab to see the progress.

![Screenshot shows the "State" value in "Stack info" tab.](/files/ElcWSxmde3BEveYyvznJ)

7\. After completing stack creation, the `Role ARN` is available from the "Outputs". Copy and save the `RoleArn` for later use.

![Screenshot shows where to obtain a Role ARN in the Outputs tab](/files/CFVRYaLhZWvbLmLzWGYE)

### Step 11. Confirm authorization

1\. Return to Select Star. You should see a form that allows you to provide "Role ARN" . Fill form in the required information:

* **AWS region**: Identifier of AWS Region where your MySQL instance is hosted eg. `us-east-2`
* **Role ARN:** Identifier of AWS IAM Role to use by Select Star.
* **Log group name prefix:** The prefix of name of the log group to be consumed by Select Star. Example: `/aws/rds/instance/dev-mysql8/general`

2\. Click **Connect**.


# Oracle (beta)

Follow these steps to connect your Oracle instance to Select Star.

1. [Create a Service Account](#id-1.-create-a-service-account)
2. [Grant additional permissions for preview access. (optional)](#id-2.-grant-additional-permissions-for-preview-access.-optional)
3. [Create a new datasource in Select Star](#id-3.-create-a-new-data-source-in-select-star)
4. [Enable query logging](#id-4.-enable-query-logging)
5. [Choose databases and schemas](#id-5.-choose-databases-and-schemas)

For instances operating within a non-publicly accessible environment, such as a private network, please refer to our guide on [Integrating Private Network Data Sources](/integrations/private-network) for detailed instructions and best practices.

## 1. Create a Service Account

We recommend creating a dedicated service account for Select Star with metadata-only roles and permissions.

Execute the following command:

```sql
-- Create the user
CREATE USER <username> IDENTIFIED BY <password>;

-- Create the role;
CREATE ROLE <role>;

-- Grant the role to the user;
GRANT <role> TO <username>;

-- Grant privileges to the role
GRANT CREATE SESSION TO <role>;
GRANT SELECT_CATALOG_ROLE TO <role>;
```

## 2. Grant additional permissions for preview access. (optional)

{% hint style="info" %}
This step is required if you wish to enable Select Star's **Preview** Feature.
{% endhint %}

Execute the following command:

```sql
-- Grant access to all tables
GRANT SELECT ANY TABLE TO <role>;
```

## 3. Create a new data source in Select Star

* **Display Name** - The Name you want to give to your new datasource.
* **Hostname** - Your hostname defines the location where your database is hosted.
* **Port** - The communication endpoint used to connect clients to the Oracle instance.
* **Username** - Specify the User to connect to the Oracle instance. It should have enough privileges to read all the metadata.
* **Password** - Password.
* **SID** - A unique identifier for an Oracle database instance (System Identifier).

![Create Data Source](/files/Bdx5Ba39MKCA0ROoOJAI)

## 4. Enable query logging

Select Star platform supports various query log sources, for optimal selection, please refer to [the documentation](/integrations/oracle/query-log).

## 5. Choose databases and schemas

After you fill in the information, you'll be asked to select the databases you'd like to load into Select Star.

{% hint style="info" %}
Select Star will not read queries or metadata for databases, schemas, or tables that are not loaded. Please load all data for which you expect to see metadata.

You can [change the databases and schemas](/data-source-management/manage-data-sources#configure-a-data-source) you have loaded if needed.
{% endhint %}

Your metadata should start loading automatically, once the sync is complete, you'll be able to see it in Select Star.

See the link below for more information on how to navigate through data warehouses.

{% content-ref url="/pages/Upyz3faFJpQ1KvNypm1x" %}
[Getting Started: Data Warehouse](/learning-data/getting-started-data-warehouse)
{% endcontent-ref %}


# Query Logs

Available query logs sources for Oracle data sources:

* System tables / views
* Database Activity Streams on an Amazon RDS instance

{% hint style="info" %}
If you are using a different hosting provider and you want to ingest query logs, please contact our sales team.
{% endhint %}

## Database Activity Streams on an Amazon RDS instance

To enable Database Activity Streams on an Amazon RDS instance running Oracle, follow the step-by-step instructions below using the AWS Management Console.

### Step 1: Create a Database Audit Specification

Execute query to create a database audit specification called `select_star_query_logs`, ensuring that all matched events are sent to the Kinesis DAS. Note that the policy is not applied to the SYS, RDSSEC, or RDSADMIN user. These are the user accounts that RDS uses to maintain the database and are not used to execute queries that Select Star care about.

The query:

```sql
CREATE AUDIT POLICY select_star_query_logs
ACTIONS ALL
WHEN 'SYS_CONTEXT(''USERENV'', ''SESSION_USER'') NOT IN (''SYS'', ''RDSSEC'', ''RDSADMIN'')'
EVALUATE PER SESSION;
```

### Step 2: Sign in to the AWS Management Console

1. Open the AWS Management Console at <https://console.aws.amazon.com/rds/>.
2. Sign in with your AWS credentials.

### Step 3: Navigate to the RDS Dashboard

1. In the AWS Management Console, search for **RDS** in the top search bar.
2. Click on **RDS** to open the Amazon RDS Dashboard.

### Step 4: Select Your RDS Instance

1. In the RDS Dashboard, click on **Databases** in the left-hand menu.
2. Locate your Oracle RDS instance in the list of databases.
3. Click on the **DB Identifier** of the Oracle instance you wish to enable the Database Activity Stream for.
4. Verify DB instance class used by instance with [supported DB instance classes for database activity streams](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/DBActivityStreams.html#DBActivityStreams.Overview.requirements.classes)

### Step 5: Start the Database Activity Stream

1. On the database details page, click on the **Actions** dropdown menu located in the top-right corner.
2. Select **Start activity stream** from the dropdown menu.

### Step 6: Configure the Database Activity Stream

1. The **Start database activity stream: name** window will appear, where `name` is your RDS instance.
2. Configure the following settings:
   * **AWS KMS key**: Choose a KMS key from the list of available AWS KMS keys. This key is used to encrypt the key that encrypts the database activity stream data.
   * **Database activity events**: Check the box to **Enable engine-native audit fields**.
3. Choose **Immediately** if you can restart the RDS instance right away and start the database activity stream immediately. If you select **During the next maintenance window**, the RDS instance will restart during the next scheduled maintenance window, and the activity stream will start at that time.

You can refer to AWS documentation about [starting a database activity stream](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/DBActivityStreams.Enabling.html) for details.

Entries in the activity stream are encrypted using unique batch keys before being written to Kineses. To facilitate decryption, the session keys is itself encrypted using the selected KMS key and included in the batch. Anyone with access to the KMS key will be able to decrypt the batch keys and use those to access raw audit events from the database.

### Step 7: Start the Activity Stream

1. After configuring the settings, click on **Start database activity stream**.
2. The status of the database will indicate that the activity stream is starting.

### Step 8: Locate AWS KMS key and AWS Kinesis stream

1. Once on the database details page, open the **Configuration** tab.
2. In the **Configuration** tab, locate in section **Database activity stream** following:
   * **AWS KMS key**: This shows the current KMS key used by the Database Activity Stream. Note this key may be different one that one used for the RDS instance.
   * **Kinesis data stream**: Note the configured Kinesis stream.

### Step 9. Create CloudFormation stack

Select Star recommends setting up integration using AWS CloudFormation, which allows you to make necessary changes to the RDS instance environment automatically, transparently, safely, and auditably. It deploys AWS Firehose and AWS Lambda function to store relevant information in AWS S3 bucket and enable us to access the AWS S3 bucket.

The source code of the CloudFormation template, build scripts, and real-time logs of the continuous deployment system are available on our public repository [selectstar/cloudformation-templates](https://github.com/selectstar/cloudformation-templates/tree/main/das), hosted on Github, to be freely audited.

You can pass the link to CloudFormation to the infrastructure team to enable the integration to be created.

1\. A simple form will be displayed in Select Star. In option **Source Type** select "RDS Database Activity Stream".

2\. Click **Open CloudFormation**. A new window will open to proceed to the creation of a CloudFormation stack by AWS Management Console. Make sure you are logged into the AWS account in which the RDS instance is hosted.

3\. The **Create Stack** form will be displayed. Some of the values will be filled in by default, under *Parameters*, enter:

* **KinesisStreamARN**: The ARN of the Kinesis stream where RDS Database Activity Streams are delivered. Example: `arn:aws:kinesis:us-east-2:000000000:stream/aws-rds-das-db-ZQO7M43PGGUXJEZVYSALTO76KA`
* **KmsKeyARN**: The ARN of the KMS key used for encryption of RDS Database Activity Streams. Example: `arn:aws:kms:us-east-2:000000000:key/f319545f-a0d4-4bfc-896f-5d37fe921ffb`
* **RdsResourceId**: The ARN of the RDS instance or cluster producing the logs. Example: `db-ZQO7M43PGGUXJEZVYSALTO76KA`

4\. Review the information and under **Capabilities** choose "I acknowledge that AWS CloudFormation might create IAM resources".

![Screenshot shows the "Capabilities" section in "Quick create stack" form.](/files/jCCAaTjHQ6eD4FWu984f)

5\. Choose **Create stack**.

6. Wait until the stack changes its status to "CREATE\_COMPLETE" from "CREATE\_IN\_PROGRESS " in the tab "Stack info." The operation should take up to 5 minutes. You need to refresh the tab to see the progress.

![Screenshot shows the "State" value in "Stack info" tab.](/files/ElcWSxmde3BEveYyvznJ)

7\. After completing stack creation, the `Role ARN` and `Bucket name` is available from the "Outputs". Copy and save the `RoleArn` and `Bucket name` for later use.

8\. Logs are published in batches, which can be buffered up to 15 minutes. After configuring and executing a database query, new files should appear in the bucket under the `processed` prefix.

### Step 10. Confirm authorization

1\. Return to Select Star. You should see a form that allows you to provide "Role ARN" . Fill form in the required information:

* **AWS region**: Identifier of AWS Region where your Oracle instance is hosted eg. `us-east-2`
* **Role ARN:** Identifier of AWS IAM Role to use by Select Star.
* **S3 bucket:** The name of the AWS S3 bucket which stores query logs. Available in the "Outputs" tab of AWS CloudFormation.
* **S3 object prefix:** The name of the AWS S3 prefix which stores query logs. For deployment using AWS CloudFormation, this is `processed/` by default.

2\. Click **Connect**.


# Salesforce (beta)

Follow these steps to connect your Salesforce instance to Select Star.

1. [Create a service account for Select Star](#id-1.-create-a-service-account-for-select-star)
2. [Required permissions](#id-2.-required-permissions)
3. [Get your security token](#id-3.-get-your-security-token)
4. [Connect Salesforce to Select Star](#id-4.-connect-salesforce-to-select-star)

For instances operating within a non-publicly accessible environment, such as a private network, please refer to our guide on [Integrating Private Network Data Sources](/integrations/private-network) for detailed instructions and best practices.

## 1. Create a service account for Select Star

We recommend creating a dedicated service account for Select Star with metadata-only roles and permissions. This account must be a **normal user with direct login credentials** and not an SSO-based user. Using an SSO-based user (such as one managed by Okta) can cause authentication issues, as the API access requires a dedicated username and password that is not subject to external identity provider restrictions.

* Go to Settings → Setup → Users → Users and click **New User**
  * **Last Name** selectstar
  * **Email** Your email
  * **User License** Salesforce Integration
  * **Profile** Minimum access - API Only Integrations *(note: if you select the minimum access profile, an administrator will need to obtain the Security Token for you)*
* Click the **Save** button
* Click **Continue**

## 2. Required permissions

To fetch metadata from Salesforce, you need the following permissions:

* **API Access** – Required to allow API access in your Salesforce organization.
* **Object Permissions** – Required to read the Salesforce objects being ingested.
* **View Setup and Configuration** – Required to fetch metadata such as Field Definitions.

## 3. Get your security token

The Salesforce Security Token is required to access the metadata through APIs.

{% hint style="info" %}
Ensure that the account used for API access is a **normal user with direct login credentials** and not an SSO-based user.
{% endhint %}

* Login to an account with administrator privileges.
* Go to Settings → Setup → Users → Users
* Locate the API-only user and click the **Login** button
* Click on the username and select **My Settings**
* From the Personal section, select **Reset My Security Token**
* Click the **Reset Security Token**
* The new token will be sent to the email address of the API-only user.

If there is no option to **login** next to the **edit** button, you need to enable the “Administrators Can Log in as Any User” option in the Setup menu:

* Go to Settings → Setup → Login Access Policies
* Select the **Administrators Can Log in as Any User** checkbox.
* Click Save.

If you need further assistance please check this [doc](https://help.salesforce.com/s/articleView?id=sf.user_security_token.htm\&type=5) on how to get the security token.

## 4. Connect Salesforce to Select Star

* **Display Name** - The Name you want to give to your new data source
* **Username** - The user's username
* **Password** - Your password
* **Security Token** - Your security token
* **Salesforce API Version** - You can specify the API version if needed, follow the steps mentioned [here](https://help.salesforce.com/s/articleView?id=000386929\&type=1) to get the API version *(default `60.0`)*
* **Salesforce Domain** - You can specify the domain you want to use to access the platform *(default `login`)*

![Create Data Source](/files/cCDVNGn6vwrVZbOXLG6j)

{% hint style="info" %}
Select Star currently does not support Popularity for this integration.
{% endhint %}

Your metadata should start loading automatically, once the sync is complete, you'll be able to see it in Select Star.

See the link below for more information on how to navigate through data warehouses.

{% content-ref url="/pages/Upyz3faFJpQ1KvNypm1x" %}
[Getting Started: Data Warehouse](/learning-data/getting-started-data-warehouse)
{% endcontent-ref %}


# Salesforce Analytics (beta)

Follow these steps to connect your Salesforce Lightning and/or CRMA instance to Select Star.

1. [Create a service account for Select Star](#id-1.-create-a-service-account-for-select-star)
2. [Required permissions](#id-2.-required-permissions)
3. [Get your security token](#id-3.-get-your-security-token)
4. [Connect Salesforce Analytics to Select Star](#id-4.-connect-salesforce-crma-to-select-star)

## 1. Create a service account for Select Star

We recommend creating a dedicated service account for Select Star with metadata-only roles and permissions. This account must be a **normal user with direct login credentials** and not an SSO-based user. Using an SSO-based user (such as one managed by Okta) can cause authentication issues, as the API access requires a dedicated username and password that is not subject to external identity provider restrictions.

* Go to Settings → Setup → Users → Users and click **New User**
  * **Last Name** selectstar
  * **Email** Your email
  * **User License** Salesforce Integration
  * **Profile** Minimum access - API Only Integrations *(note: if you select the minimum access profile, an administrator will need to obtain the Security Token for you)*
* Click the **Save** button
* Click **Continue**

## 2. Required permissions

To fetch metadata from Salesforce, you need the following permissions:

* **API Access** – Required to allow API access in your Salesforce organization.
* **Object Permissions** – Required to read the Salesforce objects being ingested.
* **View Setup and Configuration** – Required to fetch metadata such as Field Definitions.
* **Enable CRMA (CRMA)** - Permission to use CRM Analytics platform
* **Query Salesforce Objects (Lightning)** - Able to SOQL from **Report**, **Dashboard** and **DashboardComponent** (Records and Fields)

## 3. Get your security token

The Salesforce Security Token is required to access the metadata through APIs.

{% hint style="info" %}
Ensure that the account used for API access is a **normal user with direct login credentials** and not an SSO-based user.
{% endhint %}

* Login to an account with administrator privileges.
* Go to Settings → Setup → Users → Users
* Locate the API-only user and click the **Login** button
* Click on the username and select **My Settings**
* From the Personal section, select **Reset My Security Token**
* Click the **Reset Security Token**
* The new token will be sent to the email address of the API-only user.

If there is no option to **login** next to the **edit** button, you need to enable the “Administrators Can Log in as Any User” option in the Setup menu:

* Go to Settings → Setup → Login Access Policies
* Select the **Administrators Can Log in as Any User** checkbox.
* Click Save.

If you need further assistance please check this [doc](https://help.salesforce.com/s/articleView?id=sf.user_security_token.htm\&type=5) on how to get the security token.

## 4. Connect Salesforce Analytics to Select Star

* **Display Name** - The Name you want to give to your new data source
* **Integration Options** - Choose between Lightning, CRMA or both
* **Username** - The user's username
* **Password** - Your password
* **Security Token** - Your security token
* **Salesforce API Version (optional)** - You can specify the API version if needed, follow the steps mentioned [here](https://help.salesforce.com/s/articleView?id=000386929\&type=1) to get the API version *(default `64.0`)*
* **Salesforce Domain** **(optional)** - You can specify the domain you want to use to access the platform *(default `login`)*

![Create Data Source](/files/1k2jNrYZiFbT379co8XK)

{% hint style="info" %}
Select Star currently does not support Popularity for this integration.
{% endhint %}

Your metadata should start loading automatically, once the sync is complete, you'll be able to see it in Select Star.


# DB2 (beta)

Follow these steps to connect your DB2 instance to Select Star.

1. [Create a Service Account](#id-1.-create-a-service-account)
2. [Create a new datasource in Select Star](#id-2.-create-a-new-data-source-in-select-star)
3. [Choose databases and schemas](#id-3.-choose-databases-and-schemas)

For instances operating within a non-publicly accessible environment, such as a private network, please refer to our guide on [Integrating Private Network Data Sources](/integrations/private-network) for detailed instructions and best practices.

## 1. Create a Service Account

We recommend creating a dedicated service account for Select Star with metadata-only roles and permissions. To create a new DB2 user please follow the guidelines mentioned [here](https://www.ibm.com/docs/en/samfess/8.2.0?topic=schema-creating-users-manually).

DB2 user must have SELECT privilege on:

* `SYSCAT.SCHEMATA`
* `SYSCAT.TABLES`
* `SYSCAT.INDEXES`
* `SYSCAT.TABCONST`
* `SYSCAT.KEYCOLUSE`
* `SYSCAT.SQLFOREIGNKEYS`
* `SYSCAT.COLUMNS`
* `SYSCAT.VIEWS`
* `SYSCAT.SEQUENCES`

Sample SQL query to grant `SELECT` privilege on the required tables:

```sql
GRANT SELECT ON SYSCAT.SCHEMATA TO <username>;
GRANT SELECT ON SYSCAT.TABLES TO <username>;
GRANT SELECT ON SYSCAT.INDEXES TO <username>;
GRANT SELECT ON SYSCAT.TABCONST TO <username>;
GRANT SELECT ON SYSCAT.KEYCOLUSE TO <username>;
GRANT SELECT ON SYSCAT.SQLFOREIGNKEYS TO <username>;
GRANT SELECT ON SYSCAT.COLUMNS TO <username>;
GRANT SELECT ON SYSCAT.VIEWS TO <username>;
GRANT SELECT ON SYSCAT.SEQUENCES TO <username>;
```

## 2. Create a new data source in Select Star

* **Display Name** - The Name you want to give to your new datasource.
* **Hostname** - Your hostname defines the location where your database is hosted.
* **Port** - The communication endpoint used to connect clients to the DB2 instance.
* **Username** - Specify the User to connect to the DB2 instance. It should have enough privileges to read all the metadata.
* **Password** - Password.
* **Database** - The database name.

![Create Data Source](/files/seHsYiOP8BRMRlRmCilj)

{% hint style="info" %}
Select Star currently does not support Popularity and Lineage is only supported for views in this integration. If needed, please feel free to contact us so we can explore alternative solutions.
{% endhint %}

## 3. Choose databases and schemas

After you fill in the information, you'll be asked to select the databases you'd like to load into Select Star.

{% hint style="info" %}
Select Star will not read metadata for databases, schemas, or tables that are not loaded. Please load all data for which you expect to see metadata.

You can [change the databases and schemas](/data-source-management/manage-data-sources#configure-a-data-source) you have loaded if needed.
{% endhint %}

Your metadata should start loading automatically, once the sync is complete, you'll be able to see it in Select Star.

See the link below for more information on how to navigate through data warehouses.

{% content-ref url="/pages/Upyz3faFJpQ1KvNypm1x" %}
[Getting Started: Data Warehouse](/learning-data/getting-started-data-warehouse)
{% endcontent-ref %}


# PostgreSQL

## Supported Hosting Solutions

{% content-ref url="/pages/wkxBODsytNQmho2WXNJi" %}
[AWS RDS PostgreSQL](/integrations/postgres/aws-rds)
{% endcontent-ref %}

{% content-ref url="/pages/xuathWsrqsxdADsHLylN" %}
[AWS Aurora PostgreSQL](/integrations/postgres/aws-aurora)
{% endcontent-ref %}

{% content-ref url="/pages/pLZPrCl56365Yu4asdOh" %}
[PostgreSQL on-prem](/integrations/postgres/on-prem)
{% endcontent-ref %}

## Additional Information

#### Sync New Data Assets

If you create a new PostgreSQL schema in your database, we need you to re-grant our permissions.

To sync your new assets into Select Star, follow these steps:

1\. GRANT us access to the new Schema:

```sql
GRANT USAGE ON SCHEMA <schema_name> TO selectstar;
ALTER DEFAULT PRIVILEGES GRANT USAGE ON SCHEMAS TO selectstar;
```

2\. Go to **Select Star > Settings > Data**

3\. Find your Data Source, and click on **Configure**

4\. Click **Refresh**

5\. Click Next, and make sure your new Schema is selected.

![](https://github.com/selectstar/docs-customer/blob/main/integrations/.gitbook/assets/image%20\(7\)%20\(2\).png) ![](https://github.com/selectstar/docs-customer/blob/main/integrations/.gitbook/assets/image%20\(1\)%20\(1\)%20\(3\)%20\(1\).png)

We are always happy to help if you have any other questions! [Send us an email](mailto:support@getselectstar.com).


# AWS Aurora PostgreSQL

Follow these steps to connect your AWS Aurora PostgreSQL to Select Star.

## Before you start

To connect AWS Aurora to Select Star, you will need...

* access to AWS User with with permissions to deploy CloudFormation, modify AWS IAM and AWS Aurora
* access to admin user of AWS Aurora

{% hint style="info" %}
Select Star requires only minimal metadata access to AWS Aurora. The granted permissions are defined in CloudFormation template:

* IAM permission defined by resource "CrossAccountRolePolicy" in file [SelectStarAuroraPostgreSQL.json](https://github.com/selectstar/cloudformation-templates/blob/main/aurora-postgresql/SelectStarAuroraPostgreSQL.json)
* SQL command listed below

For instances operating within a non-publicly accessible environment, such as an AWS VPC private subnet, please refer to our guide on [Integrating Private Network Data Sources](/integrations/private-network) for detailed instructions and best practices.
{% endhint %}

## 1. Create PostgreSQL user

Connect to the PostgreSQL database using an administrative user account and create a new user (service account) for the integration, by executing SQL query:

```sql
CREATE USER selectstar WITH encrypted password 's313ctst8r'
```

Replace `s313ctst8r` with strong and secure password.

Then, it is necessary to grant permissions to selected databases and schemas. To do this, run the following query individually in the context of the selected database for each schema:

```sql
GRANT USAGE ON SCHEMA <schema_name> TO selectstar;
ALTER DEFAULT PRIVILEGES GRANT USAGE ON SCHEMAS TO selectstar;

```

Replace `<schema_name>` with the expected schema name and repeat for each databases & schemas.

## 2. Ensure network connectivity

To establish a connection between Select Star and your Aurora cluster, it is essential that your Aurora instance is accessible from the following IP addresses:

* 3.23.108.85
* 3.20.56.105

If your Aurora cluster is protected by a firewall, you'll need to add these two IP addresses to your whitelist to allow for the connection. If you encounter any challenges or require further assistance in adapting these configurations to your specific network topology, please don't hesitate to reach out to our technical support team for expert guidance and solutions.

## 3. Enable query logging

To be able to generate lineage and popularity, we need to have access to a log of all queries performed on the instance via AWS CloudWatch. Enabling query logging for an Amazon RDS database and sending the logs to AWS CloudWatch involves several steps. Here's a step-by-step guide to achieve this using the AWS Management Console:

1\. **Sign in to AWS Console:** Log in to your AWS Management Console using your credentials.

2\. **Open RDS Dashboard:** Navigate to the Amazon RDS service in the AWS Management Console.

3\. **Select the Aurora Cluster:** Choose the Amazon Aurora database cluster for which you want to enable query logging.

4\. **Modify the DB Cluster:** In the cluster details page, click the "Modify" button to make changes to the cluster configuration.

5\. **Enable the Query Logging Parameter:** In the "Modify DB Cluster" page, find the "Log exports" section. Look for the "PostgreSQL log" parameter. Set this parameter to "enabled."

6\. **Apply the Changes:** Scroll down to the bottom of the "Modify DB Cluster" page and click "Continue."

7\. **Review and Apply Changes:** Review the changes you're about to make and click "Apply immediately" if you want the changes to take effect immediately. Otherwise, choose a maintenance window for applying the changes. Click "Continue."

8\. **Create a New Parameter Group:** In the RDS dashboard, click on "Parameter groups" on the left-hand navigation pane.

9\. **Create a New Parameter Group:**

* In the RDS dashboard, click on "Parameter groups" in the left-hand navigation pane.
* Click the "Create parameter group" button.
* Provide a name for the new parameter group, e.g., "CustomAuroraParameterGroup."
* In the "Family" dropdown, select the appropriate DB engine family. For Aurora, you can choose "aurora-postgresql".
* Provide a description for the parameter group (optional).
* Click the "Create" button to create the new parameter group.

10\. **Edit the Parameter Group:**

* In the parameter group list, find your newly created parameter group, "CustomAuroraParameterGroup" and click on its name.
* In the "Parameter group details" page, find the "Parameters" tab.
* Click the "Edit parameters" button.

11\. **Set log\_min\_duration\_statement and log\_statement parameters:**

* In the "Modifiable parameters" page, you can search for parameters. In the search box, type "log\_min\_duration\_statement" and "log\_statement" one by one.
* For parameter `log_min_duration_statement` set value to `0` (to log all statements, regardless of duration).
* For parameter `log_statement` set value to `all` (to log all SQL statements).
* After setting these parameters, click the "Save changes" button.

12\. **Modify the Aurora Instance and Associate the Parameter Group:**

* In the RDS dashboard, select your Aurora DB instance (not the DB cluster).
* Click the "Modify" button for the instance.
* In the "DB parameter group" section, select the custom parameter group you created, "CustomAuroraParameterGroup," from the dropdown.
* Click "Continue" to proceed with the modification.
* Review the changes and click "Apply immediately" or select a maintenance window for the change to take effect. Then click "Continue."

13\. **Monitor the Update:** The changes will be applied to your Aurora instance. You can monitor the progress on the "Databases" page in the RDS dashboard.

14\. **Verify Query Logging:** After the changes have been applied, query logging will be enabled for your Aurora instance, and the logs will be sent to CloudWatch. You can access these logs by navigating to the CloudWatch Logs section of the AWS Management Console. Before accepting credentials, we verify whether the query log has been configured, so it is important that some queries have already been logged.

## 4. Create a new data source

1\. Go to the Select Star **Settings**. Click **Data** in the sidebar, then **+ Add** to create a new Data Source.

![Animations shows where to click to create a new data source.](/files/IP5pzA0NjiBfKRqyS8wu)

2\. Fill form in the required information:

* **Source Type:** Select "PostgreSQL"
* **Display Name:** This value is `PostgreSQL` by default, but you can override it if desired.
* **Hostname**: The public hostname of your instance.
* **Port**: The port used to connect. By default is `5432`, but you can adjust it if required.
* **Username**: The PostgreSQL user name to connect. In our examples of SQL queries above, we use `selectstar`.
* **Password**: The password of the PostgreSQL user.
* **DB Name**: Any database we have access to that is used to initiate a first connection.

![Form to create a new PostgreSQL instance](/files/779Q052xkZ0jFD0r1h4H)

3\. Click **Connect**.

## 5. Create CloudFormation stack

Select Star recommends use of AWS CloudFormation to setup integration, which allows you to make necessary changes to the AWS IAM in a automatic, transparent, safe and auditable manner.

The source code of the CloudFormation template along with build scripts and real-time logs of the continuous deployment system is available on our public repository [selectstar/cloudformation-templates](https://github.com/selectstar/cloudformation-templates/tree/main/aurora-postgresql) hosted on Github, to be freely audited.

You can pass the link to CloudFormation to the infrastructure team to enable the integration to be created.

1\. A simple form will be displayed in Select Star. In option **Source Type** select "AWS Aurora PostgreSQL".

![Form to connect query log for AWS Aurora](/files/hGUsDO04eZfylbigYSSh)

2\. Click **Open CloudFormation**. A new window will open to proceed to the creation of a CloudFormation stack by AWS Management Console. Make sure you are logged into the AWS account in which the Amazon Aurora cluster is hosted.

3\. The **Create Stack** form will be displayed. Some of the values will be filled in by default, unnder *Parameters*, enter:

* **Aurora cluster name**: Use the Amazon Aurora cluster name

4\. Review the information and under **Capabilities** choose "I acknowledge that AWS CloudFormation might create IAM resources".

![Screenshot shows the "Capabilities" section in "Quick create stack" form.](/files/jCCAaTjHQ6eD4FWu984f)

5\. Choose **Create stack**.

6\. Wait until the stack changes it status to "CREATE\_COMPLETE" from "CREATE\_IN\_PROGRESS" in tab "Stack info". The operation should take up to 5 minutes. You need to refresh tab to see the progress.

![Screenshot shows the "State" value in "Stack info" tab.](/files/ElcWSxmde3BEveYyvznJ)

7\. After completing stack creation, the `Role ARN` is available from the "Outputs". Copy and save the `RoleArn` for later use.

![Screenshot shows where to obtain a Role ARN in the Outputs tab](/files/7pyrZdHjypSPRbONiJss)

## 6. Confirm authorization

1\. Return to Select Star. You should see a form that allows you to provide "Role ARN" . Fill form in the required information:

* **Role ARN:** Identifier of AWS IAM Role to use by Select Star.
* **Aurora cluster name**: Name of the AWS Aurora cluster.

![Form to connect query log for AWS Aurora](/files/hGUsDO04eZfylbigYSSh)

2\. Click **Connect**.

## 7. Choose databases and schemas

After you fill in the information, you'll be asked to select the databases you'd like to load into Select Star.

{% hint style="info" %}
Select Star will not read queries or metadata or generate lineage for databases, schemas, or tables that are not loaded. Please load all data for which you expect to see lineage.

You can [change the databases and schemas](/data-source-management/manage-data-sources#configure-a-data-source) you have loaded if needed.
{% endhint %}

Select the database and click **Next**.

![](/files/-MkdC0euUkCjrubKDgis)

For each database you selected, you'll be able to select the schemas.

![](/files/-MkdC0evROqyuF58x_bx)

Your metadata should start loading automatically. Please allow 24-48 hours to completely generate popularity and lineage.

When the sync is complete, you'll be able to explore PostgreSQL in Select Star.

## Optional steps

The AWS Aurora logging setup may in some cases result in sensitive information being recorded and stored in Amazon CloudWatch. To assess impact, risk, and take precaution and remediation steps corresponding to your organization's workload characteristics, refer to relevant [section of AWS RDS service documentation](https://aws.amazon.com/premiumsupport/knowledge-center/rds-postgresql-cleartext-logging/) and AWS support if required.


# AWS RDS PostgreSQL

Follow these steps to connect your AWS RDS PostgreSQL to Select Star.

## Before you start

To connect AWS RDS PostgreSQL to Select Star, you will need...

* access to AWS User with with permissions to deploy CloudFormation, modify IAM and AWS RDS cluster
* access to admin user of AWS RDS

{% hint style="info" %}
Select Star requires only minimal metadata access to AWS RDS. The granted permissions are defined in CloudFormation template:

* IAM permission defined by resource "CrossAccountRolePolicy" in file [SelectStarAuroraPostgreSQL.json](https://github.com/selectstar/cloudformation-templates/blob/main/aurora-postgresql/SelectStarAuroraPostgreSQL.json)
* SQL command listed below

For instances operating within a non-publicly accessible environment, such as an AWS VPC private subnet, please refer to our guide on [Integrating Private Network Data Sources](/integrations/private-network) for detailed instructions and best practices.
{% endhint %}

## 1. Create PostgreSQL user

Connect to the PostgreSQL database using an administrative user account and create a new user (service account) for the integration, by executing SQL query:

```sql
CREATE USER selectstar WITH encrypted password 's313ctst8r'
```

Replace `s313ctst8r` with strong and secure password.

Then, it is necessary to grant permissions to selected databases and schemas. To do this, run the following query individually in the context of the selected database for each schema:

```sql
GRANT USAGE ON SCHEMA <schema_name> TO selectstar;
ALTER DEFAULT PRIVILEGES GRANT USAGE ON SCHEMAS TO selectstar;
```

Replace `<schema_name>` with the expected schema name and repeat for each databases & schemas.

## 2. Ensure network connectivity

To establish a connection between Select Star and your RDS instance, it is essential that your RDS instance is accessible from the following IP addresses:

* 3.23.108.85
* 3.20.56.105

If your RDS instance is protected by a firewall, you'll need to add these two IP addresses to your whitelist to allow for the connection. If you encounter any challenges or require further assistance in adapting these configurations to your specific network topology, please don't hesitate to reach out to our technical support team for expert guidance and solutions.

## 3. Enable query logging

To be able to generate lineage and popularity, we need to have access to a log of all queries performed on the cluster via AWS CloudWatch. Enabling query logging for an Amazon RDS instance and sending the logs to AWS CloudWatch involves several steps. Here's a step-by-step guide to achieve this using the AWS Management Console:

1\. **Sign in to AWS Console:** Log in to your AWS Management Console using your credentials.

2\. **Open RDS Dashboard:** Navigate to the Amazon RDS service in the AWS Management Console.

3\. **Create a New Parameter Group:**

* In the RDS dashboard, click on "Parameter groups" in the left-hand navigation pane.
* Click the "Create parameter group" button.
* Provide a name for the new parameter group, e.g., "CustomRDSParameterGroup."
* In the "Family" dropdown, select the appropriate DB engine family.
* Provide a description for the parameter group (optional).
* Click the "Create" button to create the new parameter group.

4\. **Edit the Parameter Group:**

* In the parameter group list, find your newly created parameter group, "CustomRDSParameterGroup," and click on its name.
* In the "Parameter group details" page, find the "Parameters" tab.
* Click the "Edit parameters" button.

5\. **Set log\_min\_duration\_statement and log\_statement parameters:**

* In the "Modifiable parameters" page, you can search for parameters. In the search box, type "log\_min\_duration\_statement" and "log\_statement" one by one.
* For parameter `log_min_duration_statement` set value to `0` (to log all statements, regardless of duration).
* For parameter `log_statement` set value to `all` (to log all SQL statements).
* After setting these parameters, click the "Save changes" button.

6\. **Select the AWS RDS instance:** Choose the Amazon RDS instance for which you want to enable query logging.

7\. **Modify the DB instance:** In the instance details page, click the "Modify" button to make changes to the instance configuration.

8\. **Enable the Query Logging Parameter:**

* In the "Modify DB instance" page, find the "Log exports" section. Look for the "PostgreSQL log" parameter. Set this parameter to "enabled."
* In the "Database options" section, look "DB parameter group" parameter and Select the custom parameter group you created, "CustomRDSParameterGroup" from the dropdown.

9\. **Apply the Changes:** Scroll down to the bottom of the "Modify DB instance" page and click "Continue."

10\. **Review and Apply Changes:** Review the changes you're about to make and click "Apply immediately" if you want the changes to take effect immediately. Otherwise, choose a maintenance window for applying the changes. Click "Continue."

11\. **Monitor the Update:** The changes will be applied to your RDS instance. You can monitor the progress on the "Databases" page in the RDS dashboard.

12\. **Verify Query Logging:** After the changes have been applied, query logging will be enabled for your RDS instance, and the logs will be sent to CloudWatch. You can access these logs by navigating to the CloudWatch Logs section of the AWS Management Console. Before accepting credentials, we verify whether the query log has been configured, so it is important that some queries have already been logged.

## 4. Create a new data source

1\. Go to the Select Star **Settings**. Click **Data** in the sidebar, then **+ Add** to create a new Data Source.

![Animations shows where to click to create a new data source.](/files/IP5pzA0NjiBfKRqyS8wu)

2\. Fill form in the required information:

* **Source Type:** Select "PostgreSQL"
* **Display Name:** This value is `PostgreSQL` by default, but you can override it if desired.
* **Hostname**: The public hostname of your instance.
* **Port**: The port used to connect. By default is `5432`, but you can adjust it if required.
* **Username**: The PostgreSQL user name to connect. In our examples of SQL queries above, we use `selectstar`.
* **Password**: The password of the PostgreSQL user.
* **DB Name**: Any database we have access to that is used to initiate a first connection.

![Form to create a new PostgreSQL instance](/files/779Q052xkZ0jFD0r1h4H)

3\. Click **Connect**.

## 5. Create CloudFormation stack

Select Star recommends use of AWS CloudFormation to setup integration, which allows you to make necessary changes to the RDS instance environment in a automatic, transparent, safe and auditable manner.

AWS CloudFormation will create AWS resources to enable us to access AWS CloudWatch logs.

The source code of the CloudFormation template along with build scripts and real-time logs of the continuous deployment system is available on our public repository [selectstar/cloudformation-templates](https://github.com/selectstar/cloudformation-templates/tree/main/rds-for-postgresql) hosted on Github, to be freely audited.

You can pass the link to CloudFormation to the infrastructure team to enable the integration to be created.

1\. A simple form will be displayed in Select Star. In option **Source Type** select "AWS RDS PostgreSQL".

![Form to connect query log for AWS RDS](/files/ReOlf9yC9Cnth49br1Kq)

2\. Click **Open CloudFormation**. A new window will open to proceed to the creation of a CloudFormation stack by AWS Management Console. Make sure you are logged into the AWS account in which the RDS instance is hosted.

3\. The **Create Stack** form will be displayed. Some of the values will be filled in by default, under *Parameters*, enter:

* **RDS instance name**: Use the Amazon RDS instance name

4\. Review the information and under **Capabilities** choose "I acknowledge that AWS CloudFormation might create IAM resources".

![Screenshot shows the "Capabilities" section in "Quick create stack" form.](/files/jCCAaTjHQ6eD4FWu984f)

5\. Choose **Create stack**.

6\. Wait until the stack changes it status to "CREATE\_COMPLETE" from "CREATE\_IN\_PROGRESS" in tab "Stack info". The operation should take up to 5 minutes. You need to refresh tab to see the progress.

![Screenshot shows the "State" value in "Stack info" tab.](/files/ElcWSxmde3BEveYyvznJ)

7\. After completing stack creation, the `Role ARN` is available from the "Outputs". Copy and save the `RoleArn` for later use.

![Screenshot shows where to obtain a Role ARN in the Outputs tab](/files/CFVRYaLhZWvbLmLzWGYE)

## 6. Confirm authorization

1\. Return to Select Star. You should see a form that allows you to provide "Role ARN" . Fill form in the required information:

* **Role ARN:** Identifier of AWS IAM Role to use by Select Star.
* **RDS instance name:** Name of the AWS RDS instance.

![Form to connect query log for AWS RDS](/files/ReOlf9yC9Cnth49br1Kq)

2\. Click **Connect**.

## 7. Choose databases and schemas

After you fill in the information, you'll be asked to select the databases you'd like to load into Select Star.

{% hint style="info" %}
Select Star will not read queries or metadata or generate lineage for databases, schemas, or tables that are not loaded. Please load all data for which you expect to see lineage.

You can [change the databases and schemas](/data-source-management/manage-data-sources#configure-a-data-source) you have loaded if needed.
{% endhint %}

Select the database and click **Next**.

![](/files/-MkdC0euUkCjrubKDgis)

For each database you selected, you'll be able to select the schemas.

![](/files/-MkdC0evROqyuF58x_bx)

Your metadata should start loading automatically. Please allow 24-48 hours to completely generate popularity and lineage.

When the sync is complete, you'll be able to explore PostgreSQL in Select Star.

## Optional steps

The AWS RDS logging setup may in some cases result in sensitive information being recorded and stored in Amazon CloudWatch. To assess impact, risk, and take precaution and remediation steps corresponding to your organization's workload characteristics, refer to relevant [section of AWS RDS service documentation](https://aws.amazon.com/premiumsupport/knowledge-center/rds-postgresql-cleartext-logging/) and AWS support if required.


# PostgreSQL on-prem

Follow these steps to connect your on-prem PostgreSQLto Select Star.

## Before you start

To connect PostgreSQL hosted on-prem to Select Star, you will need...

* access to admin user of PostgreSQL hosted on-prem

For instances operating within a non-publicly accessible environment, such as a private network, please refer to our guide on [Integrating Private Network Data Sources](/integrations/private-network) for detailed instructions and best practices.

## 1. Create PostgreSQL user

Connect to the PostgreSQL database using an administrative user account and create a new user (service account) for the integration, by executing SQL query:

```sql
CREATE USER selectstar WITH encrypted password 's313ctst8r'
```

Replace `s313ctst8r` with strong and secure password.

Then, it is necessary to grant permissions to selected databases and schemas. To do this, run the following query individually in the context of the selected database for each schema:

```sql
GRANT USAGE ON SCHEMA <schema_name> TO selectstar;
ALTER DEFAULT PRIVILEGES GRANT USAGE ON SCHEMAS TO selectstar;
```

Replace `<schema_name>` with the expected schema name and repeat for each databases & schemas.

## 2. Ensure network connectivity

To establish a connection between Select Star and your PostgreSQL instance, it is essential that your PostgreSQL instance is accessible from the following IP addresses:

* 3.23.108.85
* 3.20.56.105

If your PostgreSQL instance is protected by a firewall, you'll need to add these two IP addresses to your whitelist to allow for the connection. If you encounter any challenges or require further assistance in adapting these configurations to your specific network topology, please don't hesitate to reach out to our technical support team for expert guidance and solutions.

## 3. Create a new data source

1\. Go to the Select Star **Settings**. Click **Data** in the sidebar, then **+ Add** to create a new Data Source.

![Animations shows where to click to create a new data source.](/files/IP5pzA0NjiBfKRqyS8wu)

2\. Fill form in the required information:

* **Source Type:** Select "PostgreSQL"
* **Display Name:** This value is `PostgreSQL` by default, but you can override it if desired.
* **Hostname**: The public hostname of your instance.
* **Port**: The port used to connect. By default is `5432`, but you can adjust it if required.
* **Username**: The PostgreSQL user name to connect. In our examples of SQL queries above, we use `selectstar`.
* **Password**: The password of the PostgreSQL user.
* **DB Name**: Any database we have access to that is used to initiate a first connection.

![Form to create a new PostgreSQL instance](/files/779Q052xkZ0jFD0r1h4H)

3\. Click **Connect**.

2\. Fill form in the required information:

* **Source Type:** Select "None"

3\. Click **Connect** again.

Having no access to query logs can result in some limitations for the Select Star function. Specifically, information on popularity and lineage, except for those that result from views, will not be available. If none of the available integrations fit your needs, please feel free to contact us so that we can explore alternative solutions for log delivery.

## 4. Choose databases and schemas

After you fill in the information, you'll be asked to select the databases you'd like to load into Select Star.

{% hint style="info" %}
Select Star will not read queries or metadata or generate lineage for databases, schemas, or tables that are not loaded. Please load all data for which you expect to see lineage.

You can [change the databases and schemas](/data-source-management/manage-data-sources#configure-a-data-source) you have loaded if needed.
{% endhint %}

Select the database and click **Next**.

![](/files/-MkdC0euUkCjrubKDgis)

For each database you selected, you'll be able to select the schemas.

![](/files/-MkdC0evROqyuF58x_bx)

Your metadata should start loading automatically. Please allow 24-48 hours to completely generate popularity and lineage.

When the sync is complete, you'll be able to explore PostgreSQL in Select Star.


# AWS Glue (beta)

Learn how to connect AWS Glue to Select Star with our step-by-step guide. Discover the benefits of integrating these two powerful platforms for better data management and analysis.

### **Before you start**

{% hint style="info" %}
Ensure AWS Glue database and tables are setup in your AWS Glue instance. For details, see [Adding an AWS Glue crawler](https://docs.aws.amazon.com/glue/latest/ug/tutorial-add-crawler.html).
{% endhint %}

To connect AWS Glue to Select Star, you will need to...

1. [Connect AWS Glue to Select Star](#1.-connect-aws-glue-to-select-star)
2. [Create an AWS IAM Role ARN using Cloudformation](#2.-create-an-aws-iam-role-arn-using-cloudformation)
3. [Confirm authorization](#3.-confirm-authorization)
4. [Choose Catalogs and Databases](#4.-choose-catalogs-and-databases)

### **1**. Connect AWS Glue to Select Star

Select AWS Glue from the Add Data Source menu and provide the

**Display Name** - This value is `AWS Glue` by default, but you can overrided.

**Region** - ID of the AWS region where the cluster was created. For example `us-east-2`,`us-west-1`, `eu-central-1`

<figure><img src="/files/c926l2v33hcKJ7swkatv" alt=""><figcaption><p>Select AWS Glue from the Source Type list</p></figcaption></figure>

### 2. **Create** an AWS IAM Role ARN using Cloudformation

Select Star recommends use of AWS CloudFormation to setup integration, which allows you to make necessary changes to the AWS Glue environment in a automatic, transparent, safe and auditable manner.

AWS CloudFormation creates an AWS IAM Role to enable access for Select Star and add it to AWS Glue cluster.

The source code of the CloudFormation template along with build scripts and real-time logs of the continuous deployment system is available on public repository on GitHub "[selectstar/cloudformation-templates](https://github.com/selectstar/cloudformation-templates/tree/main/glue)" to be freely audited.

You can pass the link to CloudFormation to the infrastructure team to enable the integration to be created.

1\. A simple form will be displayed in Select Star.

<figure><img src="/files/m9RD1a1S7v74fZBQJgXf" alt=""><figcaption><p>Provide a Role ARN</p></figcaption></figure>

2\. Select the "Open CloudFormation" button. A new window will open to proceed to the creation of a CloudFormation stack by AWS Management Console. Make sure you are logged into the AWS account in which the AWS Glue cluster exist.

3\. The **Create Stack** form will be displayed. Fill form in the required information:

<figure><img src="/files/ouAJqGVJTcEDz61yfAoL" alt=""><figcaption></figcaption></figure>

4\. Review the information and under **Capabilities** choose "I acknowledge that AWS CloudFormation might create IAM resources".

<figure><img src="/files/jYEiIwzzu7a4lVoyO84i" alt=""><figcaption></figcaption></figure>

5\. Click **Create stack**.

6\. Wait until the stack changes it status to "<mark style="color:green;">CREATE\_COMPLETE</mark>" from "CREATE\_IN\_PROGRESS" in tab "**Stack Info**". The operation should take up to 5 minutes. You need to refresh tab to see the progress.

<figure><img src="/files/6uRv2al8MsfXzaGPAhhp" alt=""><figcaption></figcaption></figure>

7\. After completing stack creation, the `Role ARN` is available from the "**Outputs**". Copy and save the `RoleArn` for later use.

<figure><img src="/files/VnrptxWk9i7L5thZmWqh" alt=""><figcaption><p>Outputs tab</p></figcaption></figure>

### 3. Confirm authorization

1\. Return to Select Star. You should see a form that allows you to provide "Role ARN". Fill form in the required information:

* **Role ARN:** Identifier of AWS IAM Role to use by Select Star. You'll see this after completing [step 2.7](#2.-create-an-aws-iam-role-arn-using-cloudformation) of the instructions.

<figure><img src="/files/lNAEoBu7MVrJS8QnSltX" alt=""><figcaption><p>Enter Role ARN from Step 2.7</p></figcaption></figure>

2\. Click **Connect**.

### 4. Choose Catalogs and Databases

After you fill in the information, you'll be asked to select the catalogs and databases you'd like to load into Select Star.

{% hint style="info" %}
Select Star will not read queries or metadata or generate lineage for databases, or tables that are not loaded. Please load all data for which you expect to see lineage.

You can [change the catalog and databases](https://docs.selectstar.com/data-source-management/manage-data-sources#configure-a-data-source) you have loaded if needed.
{% endhint %}

Select the Catalog from the list (if more than one).

<figure><img src="/files/nCVuOoO2XaVfccrSgJlo" alt=""><figcaption><p>Ingesting Catalog</p></figcaption></figure>

Select Databases from the list.

<figure><img src="/files/v341giwS5GlFGIEOsBMQ" alt=""><figcaption><p>Select Databases</p></figcaption></figure>

Click Next and your metadata should start loading automatically. Please allow 24-48 hours to completely generate lineage.

<figure><img src="/files/aDfuwKnEGqTmT1SHXFt8" alt=""><figcaption></figcaption></figure>

When the sync is complete, you'll be able to explore AWS Glue in Select Star.


# dbt

Follow these steps to connect your dbt project to Select Star.

## Before you start

To connect dbt to Select Star, you will need...

* Permission to build the dbt project

{% hint style="info" %}
Select Star won't need any permissions for dbt directly, but you will need to build the dbt project and generate build artifacts.
{% endhint %}

Complete the following steps to connect dbt to Select Star.

1. [Determine what kind of dbt connection you need](#1.-determine-what-kind-of-dbt-connection-you-need)
2. [Connect dbt to Select Star](#2.-connect-dbt-to-select-star)

{% hint style="warning" %}
Note that in order to see changes to your dbt project reflected in Select Star after the initial connection, **you must** complete either [step 3](#3.-updating-build-artifacts-automatically) or [step 4](#4.-updating-build-artifacts-manually).
{% endhint %}

## 1. Determine what kind of dbt connection you need

Select Star can integrate with dbt via dbt cloud or via build artifacts (manifest.json, catalog.json, etc).

for DBT Cloud we require...

* a service token for the metadata api
* a job id for a custom job that will generate all the metadata we need
  * see [DBT Cloud - Setup a dbt Cloud Custom Job](/integrations/dbt/dbt-cloud#dbt-cloud-custom-job) for details about how to set up the job
* a `Read Only` and `Metadata Only` permission set

## 2. Connect dbt to Select Star

{% hint style="info" %}
If you are using multiple dbt projects and have cross-project dependencies see [dbt Project Dependencies](/integrations/dbt/dbt-project-dependencies).
{% endhint %}

Go to the Select Star **Settings**. Click **Data** in the sidebar, then **+ Add** to create a new Data Source.

![](/files/IP5pzA0NjiBfKRqyS8wu)

Choose **dbt** in the Source Type dropdown and then follow the instructions for how you want to connect:

{% content-ref url="/pages/MQoeDOVEP4yrkHgXgTM1" %}
[dbt Cloud](/integrations/dbt/dbt-cloud)
{% endcontent-ref %}

{% content-ref url="/pages/6HDe0KrWKWcvNOENCtHe" %}
[dbt Core (open source)](/integrations/dbt/dbt-core-open-source)
{% endcontent-ref %}

### Features

{% content-ref url="/pages/8HMlFQ7sJYR9PpNbffDd" %}
[dbt Tags](/integrations/dbt/dbt-tags-new)
{% endcontent-ref %}

{% content-ref url="/pages/zRjQfkCGTVO7Hb0RYkv1" %}
[dbt Tests](/integrations/dbt/dbt-tests-beta)
{% endcontent-ref %}


# dbt Cloud

Follow these steps to connect your dbt Cloud project to Select Star.

## 1. Setup a dbt Cloud Custom Job <a href="#dbt-cloud-custom-job" id="dbt-cloud-custom-job"></a>

Create a new DBT Job:

![](/files/MWAQXMxnMkQMFnoLMrgM)

Set it up with the following settings:

* Select the environment
* Select `Generate Docs`
* add a command `dbt compile --full-refresh`

![](/files/dfqBslrYvwoQuLZez4Cv)

Schedule, select every day at exactly hour 4. This will run the job before we run our ingestion, so your data source is always up to date.

![](/files/IJZUjMqhiLA5ERFY6jIt)

Finally, don't forget to save!

![](/files/lk9HsGnqBjfE72aZoFUh)

You should now see a page like this:

![](/files/AedCopE8EDznP4phs1YB)

To get your job id, check the URL, you should see something like this:

> [https://cloud.getdbt.com/#/accounts/\<account-id>/projects/\<project-id>/jobs/\<job-id>/](https://cloud.getdbt.com/#/accounts/43365/projects/69225/jobs/76899/)

the job id is the last number in the URL.

{% hint style="info" %}
This job needs to run at least once successfully in order for Select Star to be able to process your dbt project.
{% endhint %}

### a. Running your new job

Click the Green `Run now` button in the top right corner

![](/files/ZzJQrx4LB7xq6F7ypIgo)

Click on the new run generated in the list below:

![](/files/P7dJT8t1wwc6HqvRQI1s)

You should see the following screen:

![](/files/cnG2bmLoDAh7iGTztqQJ)

Wait until that grey `Running` status in the top left corner changes to a green `Success` status:

![](/files/f8WKPYCekF6pGZDWamHk)

Once the run finishes successfully, you can go to Select Star to add your dbt data source.

## 2. Add dbt Cloud to Select Star

From your +Add Data Source, select `dbt Cloud` in the `dbt Type` drop down

![Select dbt Cloud in the dbt Type drop down](/files/bohkuFhPZqRu4drHns2w)

Provide your service token and job id (see [#dbt-cloud-custom-job](#dbt-cloud-custom-job "mention") for setting up your custom job and getting it's Job Id).

#### Data Source Options <a href="#dsoptions" id="dsoptions"></a>

* **Display Name:** This value is `dbt` by default, but you can override it if desired.
* **Data Source:** Select one of your existing warehouse data sources to connect to (i.e. Snowflake, BigQuery, Redshift).
  * Note: if you don't have any other data source, you'll need to select the SQL dialect according to which data warehouse you use (i.e. Snowflake, BigQuery, Redshift). This will be labelled as **Dialect** instead of **Data Source**.
* **API Key:** You may either use a ...
  * dbt [service token](https://docs.getdbt.com/docs/dbt-cloud/dbt-cloud-api/service-tokens) (preferred)
    * Permissions: we need the **Metadata Only** and **Read Only** permission for the projects that you are connecting
    * For enterprise customers of dbt, we need [**Job Viewer**](https://docs.getdbt.com/docs/collaborate/manage-access/enterprise-permissions#job-viewer) and [**Account View**](https://docs.getdbt.com/docs/collaborate/manage-access/enterprise-permissions#account-viewer) permissions.
  * dbt [user token](https://docs.getdbt.com/docs/dbt-cloud/dbt-cloud-api/user-tokens)
* **Job Id:** The id for the custom job which will generate the dbt metadata needed for ingestion
* **Access URL**: The [URL for your account ](https://docs.getdbt.com/docs/cloud/about-cloud/access-regions-ip-addresses#accessing-your-account)(given by dbt, based on your region, tenancy and plan). For most accounts this will be the default <https://cloud.getdbt.com>, which can be updated if needed.

Click the **Connect** button at the bottom of the **Add Data Source** modal.

Your metadata should start loading automatically. When the sync is complete, we'll email you and you'll be able to explore dbt in Select Star!


# dbt Core (open source)

Follow these steps to connect your dbt Core project to Select Star.

## 1. Create dbt build artifacts

For dbt Core, we'll need

* `manifest.json` file
* `catalog.json` file
* `run_results.json` for dbt Tests ingestion. `run_results.json` is optional.

To generate these, please use these commands:

for `manifest.json`, run the following command and copy the `manifest.json` in your `target` directory to another location.

```
dbt compile --full-refresh
```

for `catalog.json`, run the following command

```
dbt docs generate
```

for dbt Tests, run the following command:

```
dbt test
```

after generating these, you should see them in your `target` directory.

{% hint style="danger" %}
`dbt docs generate` also generates a `manifest.json` file, but this doesn't contain enough information about lineage, do ***not*** provide this manifest.json to Select Star if you are using **Create a separate data source**.

This is why we recommend copying the `manifest.json` to another location
{% endhint %}

## 2. Add dbt Core to Select Star

Select `dbt Core (open source)` in the `dbt Type` drop down

![](/files/WkL8L27xfRSrCSbrRcOC)

Provide the `manifest.json` and `catalog.json` and `run_results.json` files into the file drop zones. Please note that `run_results.json` is optional and only required if you want to see dbt tests in Select Star.

#### Data Source Options <a href="#dsoptions" id="dsoptions"></a>

* **Display Name:** This value is `dbt` by default, but you can override it if desired.
* **Data Source:** Select one of your existing warehouse data sources to connect to (i.e. Snowflake, BigQuery, Redshift).
  * Note: if you don't have any other data source, you'll need to select the SQL dialect according to which data warehouse you use (i.e. Snowflake, BigQuery, Redshift). This will be labelled as **Dialect** instead of **Data Source**.
* **manifest.json:** Drag and drop a `manifest.json` file from your dbt project's target directory.
* **catalog.json:** Drag and drop a `catalog.json` file from your dbt project's target directory.
* **run\_results.json:** Drag and drop a `run_results.json` file from your dbt project's target directory.

Click the **Connect** button at the bottom of the **Add Data Source** modal.

Your metadata should start loading automatically. When the sync is complete, we'll email you and you'll be able to explore dbt in Select Star!

{% hint style="warning" %}
**Important:** If you are using dbt Core, in order to see changes to dbt reflected in Select Star, you will need to update the build artifacts after making changes.

You can update the build artifacts automatically (see [step 3](#3.-updating-build-artifacts-automatically)) or manually (see [step 4](#4.-updating-build-artifacts-manually)).
{% endhint %}

## 3. Update build artifacts automatically (recommended)

{% hint style="info" %}
if you are using `dbt Cloud` then you don't need to do this, we'll automatically pull updates and the job created in [#dbt-cloud-custom-job](#dbt-cloud-custom-job "mention")will keep the metadata up to date!
{% endhint %}

In order to keep your dbt docs to be synched to Select Star, you can automatically update the `manifest.json` and `catalog.json` files using our dedicated [API endpoint](https://api.production.getselectstar.com/docs/#operation/ingestion_dbt_create).

Learn how to integrate with the Select Star API by clicking the link below.

{% content-ref url="/pages/-Mkw\_00H3MfUYEoEtgaK" %}
[Select Star API](/select-star-api)
{% endcontent-ref %}

## 4. Update build artifacts manually

Navigate to the data source admin page, and next to the dbt data source click on the **Configure** button.

![](/files/rIHhLdYWhPko7he1rZao)

Here you can select the `dbt Core (open source)` dbt Type then update the `manifest.json` and `catalog.json` files in by dragging and dropping them from your dbt project's target directory.

See the steps in [**Data Source Options**](#dsoptions-1) to see how to create your `manifest.json, catalog.json` and `run_results.json` files.

Once you've updated your `manifest.json` and/or `catalog.json and run_results.json` files, click **Connect** to confirm your changes, and you're good to go!


# dbt Tags

Learn how to access and view your dbt tags in Select Star with ease. Our simple guide will show you how to access your tags via the Tags sidebar.

We now sync all dbt tags to Select Star!

You can see your dbt tags as part of your Tags sidebar once they are ingested in Select Star.

![](/files/ACbyBMDZFdegbnIigRqn)

Tags in Select Star will automatically be matched to dbt tags ingested, based on their names.

To manually match a dbt tag to Select Star tag:

1. Go to admin and the dbt data source within the admin settings
2. Go to the Tags Tab
3. Use the drop down to match any Select Star tab to a dbt Tag.

<figure><img src="/files/JxNHXzF3LldyPAYewcJL" alt=""><figcaption><p>Match dbt tags to Select Star tags</p></figcaption></figure>

Once matched, the Select Star tag will show up everywhere instead of the dbt tag with a tooltip highlighting this connection.

<figure><img src="/files/rUNGH89w435ydwJz7QM8" alt=""><figcaption><p>Linked PII tags showing up as Select Star tag</p></figcaption></figure>

To hide a dbt Tag from Select Star:

1. Go to admin and the dbt data source within the admin settings
2. Go to the Tags Tab
3. Uncheck any dbt tags that are not needed in Select Star.

To learn more about how tags in Select Star work, go to [Tags](/features/tags)


# dbt Tests

Find out the latest news that dbt tests are now going to be ingested into Select Star, and how this will benefit your data projects.

dbt Tests are now going to be ingested into Select Star!

{% hint style="info" %}
To add dbt tests into your existing SelectStar organization, you will need to either update permissions (dbt Cloud) or add a `run_results.json` (dbt Core)
{% endhint %}

To see how to update your config:

{% content-ref url="/pages/MQoeDOVEP4yrkHgXgTM1" %}
[dbt Cloud](/integrations/dbt/dbt-cloud)
{% endcontent-ref %}

{% content-ref url="/pages/6HDe0KrWKWcvNOENCtHe" %}
[dbt Core (open source)](/integrations/dbt/dbt-core-open-source)
{% endcontent-ref %}

Once updates, to find dbt tests, go to the dbt tests page from the hierarchy

<div align="left"><figure><img src="/files/FzmLa9f4dz9wwsHiIiCy" alt="" width="292"><figcaption><p>dbt tests in hierarchy</p></figcaption></figure></div>

All tests that have run in the dbt data source should be available in the dbt tests page.

<figure><img src="/files/oHQQ8ZJkOrBXX3caSyD0" alt=""><figcaption><p>Full dbt Tests Page</p></figcaption></figure>

To filter for failed tests, click on one of the tabs on top:

<figure><img src="/files/q1Eoz191ICThzNDe84YD" alt=""><figcaption><p>Failed dbt tests</p></figcaption></figure>

To view dbt tests that have run for a particular model or table, we can go to the newly added tests tab in each table/model page:

<figure><img src="/files/OO9VcSmqHYqPbAyB53Ob" alt=""><figcaption><p>Test that have run for view model.</p></figcaption></figure>

Click on any of the tests to drill down even further to see information such as all test runs and errors from the latest test run and the last run date.

<figure><img src="/files/paKBIuRjS0o8idnU61AA" alt=""><figcaption><p>Failed test ingested from dbt</p></figcaption></figure>

You can each see the source of each test and copy this information really quickly:

<figure><img src="/files/l612RPaivUyo5cJRj2A4" alt=""><figcaption><p>Source of test not_null</p></figcaption></figure>

dbt tests are also searchable like any other asset type in Select Star

<figure><img src="/files/sTbfSWfIxSX08QvsPKX6" alt=""><figcaption><p>Search by the name of the dbt test</p></figcaption></figure>


# dbt docs Sync

With dbt docs Sync, customers can update table and column descriptions in Select Star and have them seamlessly synced back to your dbt repository.

## **Key Benefits**

1. **Automated Documentation Sync and Version Control Transparency:** Say goodbye to manual updates and let Select Star do the work! dbt docs Sync automatically transfers table and column descriptions from your Select Star documentation to your version control repository.
2. **Enhanced Collaboration:** Empower your team to collaborate seamlessly by maintaining clear and up-to-date descriptions. The shared documentation will provide valuable insights to your team members, ensuring everyone stays on the same page.
3. **Ease of use:** Documenting your dbt tables and columns is a walk in the park with Select Star, unlike the old-school YML file editing. Instead of grappling with code, Select Star offers a user-friendly interface that allows you to seamlessly add and update descriptions directly. This intuitive approach ensures that your documentation process is not only simpler but also more accessible to your team members, fostering efficient collaboration.

## Requirements

1. **Git Repository**\
   The dbt project must be hosted in a Github or Bitbucket repository
2. **Schema File**

   Select Star syncs descriptions from and to models. While you can create models by just creating a sql file, **we require that you have a `schema.yml` file in your dbt repository in order to sync the descriptions** back to your model.
3. **Validated Descriptions**

   Select Star can display descriptions for columns in many different ways: Suggested descriptions propagated up/downstream, AI generated descriptions, etc. For us to sync back descriptions to your dbt repository **we require that the description is manually confirmed or entered by the user**.

## Sync Frequency

A Pull Request (PR) is automatically created in your Git Repo at the scheduled time, which is currently set to 12pm UTC. If there are any pending changes to the descriptions, a new PR will be generated daily.

If there is an open PR created by Select Star, the daily scheduled job will close the branch, perform a Git sync, and then create a new PR. This process helps us steer clear of any merge conflicts in your repository.

{% content-ref url="/pages/m3xDhga9QdY0UJCemmk6" %}
[Github dbt docs Sync](/integrations/dbt/dbt-docs-sync/github-dbt-docs-sync)
{% endcontent-ref %}

{% content-ref url="/pages/UICikEGKl4TBoMyOBOjX" %}
[Bitbucket dbt docs Sync](/integrations/dbt/dbt-docs-sync/bitbucket-dbt-docs-sync)
{% endcontent-ref %}

{% hint style="info" %}
This feature is in Private Preview. Please reach out to <support@getselectstar.com> if you to enable this feature.
{% endhint %}


# Github dbt docs Sync

Connect your Github dbt repository directly in Select Star with the following steps

1. Using the Github Configuration UI available in the Apps section, add the Github repo details.

   <figure><img src="/files/J0fUU1cqnjXwl8xo6p0q" alt=""><figcaption></figcaption></figure>
2. To ensure Select Star can access your Github Repo, install the Github app on your repository.\
   \
   Github App: [Select Star dbt docs Sync](https://github.com/apps/select-star-dbt-docs-sync)\\

   **Github App Required Permissions**

   * **Read access to metadata:** This allows us to ensure we are syncing the descriptions to the updated models.
   * **Read and write access to code and pull request:** This allows us to create or delete a PR and verify that there are no conflicts when generating a PR.

<div data-full-width="true"><figure><img src="/files/Go937n331KezoXpHqqWg" alt=""><figcaption><p>Install the Github App on your Github Repo</p></figcaption></figure></div>

<figure><img src="/files/LkWIWSo7laPPTBIZqu3C" alt=""><figcaption><p>Github App: Select Star dbt docs Sync</p></figcaption></figure>

<figure><img src="/files/uVmtMXyKBGE8QA6bbPmd" alt=""><figcaption><p>Select the list of repositories</p></figcaption></figure>

3. Provide the following information

* Select Star Datasource - This dropdown will have a list of all your connected dbt Datasources
* Github URL - Provide the Github Repository URL
* Project Name - dbt Project name, as mentioned in the `dbt_project.yml` file, and
* Path Prefix - Information about the Root folder (if the dbt project is not in the root folder)

<figure><img src="/files/k6iTHqcjVwm6w60mQasx" alt=""><figcaption></figcaption></figure>

Once this is setup, you will see new PRs generated in your Github Repo.

<figure><img src="/files/F4zhnrCzrko729lm6mxW" alt=""><figcaption><p>Github PR</p></figcaption></figure>

You will also be able to see the list of all the PRs created by this feature in the Select Star UI.

<figure><img src="/files/BftKB9stqCGj2tsOgCTP" alt=""><figcaption></figcaption></figure>


# Bitbucket dbt docs Sync

To connect your Bitbucket dbt repository, you will need to follow these steps -

1. Create a Repository Access Token in Bitbucket and grant the following permissions -

* **Read and write access to repositories:** This allows us to ensure we are syncing the descriptions to the updated models.
* **Read and write access to pull request:** This allows us to create or delete a PR and verify that there are no conflicts when generating a PR.

<figure><img src="/files/ycimTzXMe2M45HtoRPJP" alt=""><figcaption><p>Create Repository Access Token</p></figcaption></figure>

2. Provide the following information -
   1. Bitbucket Repository URL,
   2. dbt Project name, as mentioned in the `dbt_project.yml` file, and
   3. Root folder (if the dbt project is not in the root folder).

Once this is setup, you will see new PRs generated in your Bitbucket Repo.

<figure><img src="/files/6Lf5QEqtdlAMITIfDOui" alt=""><figcaption></figcaption></figure>


# dbt Impact Report

Prevent breaking changes before they break your dashboards and key models with Select Star's dbt Impact Report.

## Introduction

Select Star's dbt Impact Report allows you to make model changes with confidence. The Impact Report shows downstream items for your dbt model changes right in your pull request, so you can view potential impact before making changes and prevent downstream breakages.

dbt Impact Report is available via GitHub Actions. Follow the instructions at <https://github.com/selectstar/dbt-impact-report-action> or below to set up the GitHub Action.

<figure><img src="/files/6JzeJ4w8XMBDBnbhhRcq" alt=""><figcaption></figcaption></figure>

## Before You Start

1. dbt MUST be **set up as a separate data source**

<div data-full-width="true"><figure><img src="/files/JG5ZyWKX8w3Sgd27IAE5" alt="" width="563"><figcaption></figcaption></figure></div>

1. **Select Star API Token** - this is required for the Action to use Select Star's APIs. See [API Token](/select-star-api/authentication).
2. **Select Star Data Source GUID** - this is the GUID of the dbt data source corresponding to the repository you're adding the Action to.
   1. You can get the GUID by going into **Admin > Data** and selecting the dbt data source. The GUID will be in the URL and look like `ds_example` .

<div data-full-width="true"><figure><img src="/files/oC19QYVbBSJiYaHrsvwY" alt=""><figcaption></figcaption></figure></div>

## Configure the GitHub Action

1. If you don't already have it, create repository secrets in your repo:
   1. `SELECTSTAR_API_TOKEN` with the value of the API Token from the prerequisites.
2. Add the Select Star GitHub Action to your workflow:
   1. Create a workflow file in your repo
   2. Add the following code to the workflow file

{% code lineNumbers="true" %}

```yaml
name: Select Star dbt impact report

on:
  pull_request:
    types: [opened, edited, synchronize, reopened]

jobs:
  create-impact-report:
    name: Create the impact report for dbt projects
    runs-on: ubuntu-latest
    permissions:
      pull-requests: write 
    steps:
      - name: Run Action
        uses: selectstar/dbt-impact-report-action@v1
        with:
          GIT_REPOSITORY_TOKEN: ${{secrets.GITHUB_TOKEN}}   # no need to change, GitHub will handle it as it is
          SELECTSTAR_API_TOKEN: ${{secrets.SELECTSTAR_API_TOKEN}}
          SELECTSTAR_API_URL: YOUR INSTANCE API URL   # (e.g.: https://api.production.selectstar.com)
          SELECTSTAR_WEB_URL: YOUR INSTANCE WEB URL   # (e.g.: https://app.selectstar.com)
          SELECTSTAR_DATASOURCE_GUID: YOUR DBT DATA SOURCE GUID  # (e.g.: ds_aRjCTzAf4dPNigiV87Uggq)
```

{% endcode %}

{% hint style="warning" %}
Make sure the Datasource's GUID you add to the Github Workflow is the one corresponding to your DBT data source. Adding other GUIDs will not yield the right results in the report.
{% endhint %}

## Test it Out

After configuring the GitHub action, test out the dbt Impact Report by creating a pull request with any change to a dbt model file in the repo. You should see the action running and a new comment generated on the pull request with the Impact report.


# dbt Project Dependencies

When using dbt projects as dependencies it is recommended to add all of them into a single datasource, this would give you a more complete lineage view while decreasing code duplication.

## How to manage your projects

Once your data source is created you will se a "Projects" tab in the data source settings where you can Add or remove additional projects.

<figure><img src="/files/0Cr8ynTkJnnowv8apQqR" alt=""><figcaption><p>dbt projects settings</p></figcaption></figure>

## Editing a data source with multiple projects

When configuring a datasource that has more than one project there will be a notification at the bottom of the form letting you know this is the case and pointing you to the projects settings.

<figure><img src="/files/vWVL9Oacc6LugfEo4BnV" alt=""><figcaption><p>configure datasource with multiple projects</p></figcaption></figure>

## How the project's data is processed

1. Once your datasource is ingested all projects will be processed, de-duplicating the objects found
2. The lineage view will now show lineage across dbt objects across all projects that have been added in the same data source


# Apache Airflow (beta)

Follow these steps to connect your Apache Airflow to Select Star (via OpenLineage).

## Before you start

To connect Apache Airflow to Select Star, you will need...

* Permission to install and update packages in your Airflow environment

{% hint style="info" %}
Select Star won't need any permissions for your Airflow directly, but you will need to install a Python package and configure an environment variable in your Airflow environment.
{% endhint %}

Complete the following steps to connect Apache Airflow to Select Star.

1. [Create a new Data Source in Select Star](#id-1.-create-a-new-data-source-in-select-star)
2. [Configure Apache Airflow](#id-2.-configure-apache-airflow)
3. [Sync Metadata in Select Star](#id-3.-sync-metadata-in-select-star)

{% hint style="info" %}
Note that Select Star does not connect to Apache Airflow directly. Instead, we connect via [OpenLineage](https://openlineage.io/), which is an open platform for collection and analysis of data lineage. It tracks metadata about your Apache Airflow datasets and DAGs, DAG Runs, and sends that metadata to Select Star.\
\
Airflow DAGs will not appear in the catalog until metadata is received and ingestion is run.
{% endhint %}

## 1. Create a new Data Source in Select Star

Go to the Select Star Settings. Click Data in the sidebar, then + Add to create a new Data Source.

<figure><img src="/files/MB2zOCxnnLBBK4E1M9yK" alt=""><figcaption></figcaption></figure>

Fill the form in the required information:

* **Display Name** - This value is `Apache Airflow` by default, but you can override it.
* **Source Type** - Choose `Apache Airflow` from the dropdown.
* **Base URL** - The URL of your Apache Airflow instance. For example, `http://airflow.example.com`.

Click **Save** to proceed.

On the next screen, you will get the **API Token**, **Events Endpoint** and the **Events URL**. You will need these in the next steps to configure your Apache Airflow environment.

<figure><img src="/files/V4S8t3LXcAbFAoBjtwO1" alt=""><figcaption></figcaption></figure>

* **API Token** - This is a secret key that Select Star will use to authenticate the traffic coming from your Apache Airflow instance.
* **Events Endpoint** - This is the Select Star endpoint where your Apache Airflow instance will send OpenLineage events, containing the metadata about your DAGs, DAG Runs, and datasets.
* **Events URL** - This is the Select Star Base URL where your Apache Airflow instance will send OpenLineage events.

## 2. Configure Apache Airflow

### Install OpenLineage provider

Install the provider package or add the following line to your requirements file *(usually requirements.txt)*:

```
apache-airflow-providers-openlineage==1.10.0
```

### Transport setup

1. Self-hosted Apache Airflow Provide a Transport configuration so that OpenLineage knows where to send the events. Keep the API Token and Events Endpoint handy from the previous step.

* Within airflow\.cfg file

```properties
[openlineage]
transport = {"type": "http", "url": "<EVENTS_URL_PROVIDED_BY_SELECT_STAR>", "endpoint": "<EVENTS_ENDPOINT_PROVIDED_BY_SELECT_STAR>", "auth": {"type": "api_key", "api_key": "<API_KEY_PROVIDED_BY_SELECT_STAR>"}}
```

* or with `AIRFLOW__OPENLINEAGE__TRANSPORT` environment variable

```bash
AIRFLOW__OPENLINEAGE__TRANSPORT='{"type": "http", "url": "<EVENTS_URL_PROVIDED_BY_SELECT_STAR>", "endpoint": "<EVENTS_ENDPOINT_PROVIDED_BY_SELECT_STAR>", "auth": {"type": "api_key", "api_key": "<API_KEY_PROVIDED_BY_SELECT_STAR>"}}'
```

{% hint style="info" %}
Make sure to replace `<EVENTS_URL_PROVIDED_BY_SELECT_STAR>`, `<EVENTS_ENDPOINT_PROVIDED_BY_SELECT_STAR>` and `<API_KEY_PROVIDED_BY_SELECT_STAR>` with the actual values provided by Select Star in [Step 1](https://github.com/selectstar/docs-customer/blob/main/integrations/apache-airflow/README.md#id-1.-create-a-new-data-source-in-select-star)
{% endhint %}

2. Amazon Managed Workflows for Apache Airflow (MWAA) In the case of Amazon MWAA, the installation of OpenLineage does not change, however, setting up transport is done using the plugin.

First, create `env_var_plugin.py` file. Paste the following code:

```python
from airflow.plugins_manager import AirflowPlugin
import os

os.environ["AIRFLOW__OPENLINEAGE__NAMESPACE"] = "airflow"
os.environ["AIRFLOW__OPENLINEAGE__TRANSPORT"] = '''{
  "type": "http", 
  "url": "<EVENTS_URL_PROVIDED_BY_SELECT_STAR>",
  "endpoint": "<EVENTS_ENDPOINT_PROVIDED_BY_SELECT_STAR>",
  "auth": { 
    "type": "api_key", 
    "api_key": "<API_KEY_PROVIDED_BY_SELECT_STAR>"
   }
}'''
os.environ["AIRFLOW__OPENLINEAGE__CONFIG_PATH"] = ""
os.environ["AIRFLOW__OPENLINEAGE__DISABLED_FOR_OPERATORS"] = ""


class EnvVarPlugin(AirflowPlugin):
    name = "env_var_plugin"
```

{% hint style="info" %}
Make sure to replace `<EVENTS_URL_PROVIDED_BY_SELECT_STAR>`, `<EVENTS_ENDPOINT_PROVIDED_BY_SELECT_STAR>` and `<API_KEY_PROVIDED_BY_SELECT_STAR>` with the actual values provided by Select Star in [Step 1](https://github.com/selectstar/docs-customer/blob/main/integrations/apache-airflow/README.md#id-1.-create-a-new-data-source-in-select-star)
{% endhint %}

If you already have `plugins.zip` file, add `env_var_plugin.py` to it. Otherwise, you can create it by calling:

```bash
zip plugins.zip env_var_plugin.py
```

Update your plugins in MWAA environment by following these steps:

* Upload `plugins.zip` to the S3 bucket associated with MWAA environment.
* Go to your MWAA environment
* Click `Edit`
* Scroll to section `DAG code in Amazon S3`
* Under `Plugins file` choose your `plugins.zip` file and set the version to the latest

*NOTE: You should do the same for `requirements.txt` file*

Now environment will update itself by downloading and installing the plugin. It may take a while for changes to take effect.

**That’s it!** OpenLineage events should be sent to the Select Star when DAGs are run.

*For more details on using OpenLineage integration with Apache Airflow, please read the* [*official airflow documentation*](https://airflow.apache.org/docs/apache-airflow-providers-openlineage/stable/guides/user.html)*.*

## 3. Sync Metadata in Select Star

After you have configured your Apache Airflow environment, make sure to trigger your healthcheck DAGs. This will send OpenLineage events to Select Star, and help you verify that the integration is working correctly.

Afterwards, you can go to the Select Star Settings and click on the Data in the sidebar. Click on the Sync metadata button on your Apache Airflow Data Source.

{% hint style="warning" %}
Note that Select Star does not connect to Apache Airflow directly. That means the lineage and your DAGs metadata will be available in Select Star only after you run your DAGs and OpenLineage events are sent to Select Star and ingestion is completed.
{% endhint %}

If you want to examine OpenLineage events without sending them anywhere, you can set up ConsoleTransport. The events will end up in task logs.

```properties
[openlineage]
transport = {"type": "console"}
```


# OpenLineage (beta)

Follow these steps to start pushing your OpenLineage events to Select Star.

## Before you start

To connect your Data Sources to Select Star, via OpenLineage, you will need to...

* prepare your OpenLineage job events according to the [specifications](https://openlineage.io/docs/spec/facets/)

{% hint style="info" %}
Select Star does not need any permissions for your underlying data sources or ETL tools, rather it relies on the events prepared and pushed by you.
{% endhint %}

Complete the following steps to connect OpenLineage to Select Star.

1. [Create a new Data Source in Select Star](#id-1.-create-a-new-data-source-in-select-star)
2. [Configure OpenLineage Producer](#id-2.-configure-openlineage-producer)
3. [Sync Metadata in Select Star](#id-3.-sync-metadata-in-select-star)

## 1. Create a new Data Source in Select Star

Go to the Select Star Settings. Click Data in the sidebar, then + Add to create a new Data Source.

<figure><img src="/files/UOcyzCP9jK7RLZfrVl09" alt=""><figcaption></figcaption></figure>

Fill in the form with the required information:

* **Source Type** - Choose `OpenLineage` from the dropdown.
* **Display Name** - This value is `OpenLineage` by default, but you can override it.
* **Base URL** - The URL of your ETL instance. For example, `http://airflow.example.com`.

Click **Save** to proceed.

On the next screen, you will see the **API Token**, **Events Endpoint**, and the **Events URL**. You will need these in the next steps to configure your OpenLineage producer environment.

<figure><img src="/files/bUFO0ydlMr3w0L4Bkmth" alt=""><figcaption></figcaption></figure>

* **API Token** - This is a secret key that Select Star will use to authenticate the traffic coming from your OpenLineage producer instance.
* **Events Endpoint** - This is the Select Star endpoint where your producer will send OpenLineage events, containing the metadata about your Jobs, Job Runs, and Datasets.
* **Events URL** - This is the Select Star Base URL where your producer will send OpenLineage events.

## 2. Configure OpenLineage Producer

You must use the values provided above in your producer to start sending your OpenLineage events to Select Star.

**That's it!** OpenLineage events will be sent to Select Star when your producer starts creating events.

{% hint style="info" %}
Select Star's OpenLineage integration can be used to generate lineage from your Spark, Airflow, and Custom OpenLineage events.
{% endhint %}

*For more details on configuring and producing OpenLineage events, please read the* [*official openlineage documentation*](https://openlineage.io/getting-started)*.*

## 3. Sync Metadata in Select Star

After you have configured your OpenLineage environment, make sure to trigger your health check jobs. This will send OpenLineage events to Select Star, and help you verify that the integration is working correctly.

Afterwards, you can go to the Select Star Settings and click on the Data in the sidebar. Click on the Sync metadata button on your OpenLineage Data Source.

{% hint style="warning" %}
Note that Select Star does not connect to your sources directly. That means the lineage and your job metadata will be available in Select Star only after you run your jobs and OpenLineage events are sent to Select Star.
{% endhint %}


# Tableau

The Select Star integration allows users to view, search and understand their Tableau instance all in one convenient place. Discover the benefits and features of the Tableau Select Star integration in

{% content-ref url="/pages/BSpSyNVrqdRyOPQG0HVa" %}
[Tableau Cloud](/integrations/tableau-server/tableau-online)
{% endcontent-ref %}

{% content-ref url="/pages/R9AkRTPfc1xEGyPFNHXV" %}
[Tableau Server](/integrations/tableau-server/tableau-server)
{% endcontent-ref %}

{% content-ref url="/pages/-MgSXxjMXQRitRJRxsSG" %}
[Getting Started: Tableau](/learning-data/getting-started-tableau)
{% endcontent-ref %}


# Tableau Cloud

Select Star, a leading data management platform, now integrates seamlessly with Tableau Online. Discover how this integration can help you optimize your data visualization and analysis capabilities.

## Requirements

* REST API and Metadata API enabled
* A valid Personal Access Token

## 1. Enable REST API and Metadata API

Tableau's [REST API](https://help.tableau.com/current/api/rest_api/en-us/REST/rest_api_requ.htm) and [Metadata API](https://help.tableau.com/current/api/metadata_api/en-us/index.html) **must** be enabled in order to see Tableau metadata in Select Star.

The REST API is permanently enabled for Tableau Cloud, and is enabled by default on all supported versions of Tableau Server.

To enable the Metadata API, a server admin must enable the Metadata API on Tableau Server using the `tsm maintenance metadata-services enable` command through the Tableau Services Manager (TSM) command line interface (CLI). For more information, see [Tableau Documentation](https://help.tableau.com/current/api/metadata_api/en-us/docs/meta_api_start.html#enable-the-tableau-metadata-api-for-tableau-server).

For both Tableau Cloud and Tableau Server, go to **Settings > Automatic Access to Metadata about Databases and Tables**, and make sure the checkbox is checked.

<figure><img src="/files/D4Sb0W88Ooj4PwEfIi91" alt=""><figcaption><p>Automatic Access to Metadata about Databases and Tables</p></figcaption></figure>

{% hint style="info" %}
**Tableau Cloud** can take some time to populate the Metadata Database if you have a big amount of Workbooks in Tableau Cloud.
{% endhint %}

## 2. Create a Personal Access Token

Navigate to **My Account Settings** on the top right of the screen.

![My Account Settings](/files/TeuqLewpc3hoklGBOsg0)

Scroll down to **Personal Access Tokens**.

Enter a name for your token and click **Create new token.**

{% hint style="info" %}
The user you connect to with Select Star needs a **Site Administrator Explorer** role at minimum in order for Select Start to retrieve all the necessary metadata from the Tableau API.
{% endhint %}

![](/files/1zHnrReBTewCpMdy7mMP)

Copy the **Token Secret** from the popup and store it somewhere safe. You will need them later on during this setup.

## 3. Create a Custom SQL View (Tableau Cloud)

{% hint style="info" %}
It is mandatory to create a Custom SQL View from Admin Insights. Without creating this Workbook, you will not be able to ingest information from Tableau.
{% endhint %}

Go to **Site Status** > **Admin Insights**.

Click on **New** > **Upload Workbook.**

![](/files/WA80p1cFOF8HdyXh7M15)

[Download this .twbx file](https://drive.google.com/file/d/1RWgZLZpW3eKIvHZVtdybK2YuNzua2ZMa/view?usp=sharing) and upload as the Tableau Workbook.

Confirm that **View Stats** & **Views** are populating correctly inside the Select Star Admin Insights workbook.

## 4. Create a Tableau Data Source

1. Go to the Select Star **Settings > Data**
2. Click **+ Add** to create a new Data Source.

Fill out the set up form with the information as outlined below.

* **Display Name:** Give your data source a meaningful name, the default is Tableau.
* **Access Token Name:** The Personal Access Token Name you created in [Step 2](#2.-create-a-personal-access-token).
* **Access Token Secret:** The Personal Access Token Secret created in [Step 2](#2.-create-a-personal-access-token).
* **Base Url:** The main URL for your site, such as `https://10ax.online.tableau.com`
* **Site Id:** The name of your site in tableau. You can extract this from the URL you access in Tableau. Example, ***analytics*** is the name of the site for the URL `https://10ax.online.tableau.com/#/site/analytics/home`.

Click the **Connect** button at the bottom of the modal.

<figure><img src="/files/iySXTJh38qD0fAVATtKY" alt=""><figcaption></figcaption></figure>


# Tableau Server

Learn how Select Star integrates with Tableau Server to improve data access and visibility, and how this integration can help your organization make more informed decisions based on real-time data.

## Requirements

* REST API and Metadata API enabled
* A valid Personal Access Token
* [Tableau Server version 2020.4.6](https://www.tableau.com/support/releases/server/2020.4.6#esdalt) (build number 20204.21.0618.0842) or later. If you need to upgrade your Tableau Server, here are a few helpful links:
  * [Preparing for Upgrade](https://help.tableau.com/current/server-linux/en-us/server-upgrade-prepare.htm)
  * [Upgrading from 2018.1 and Later (Linux)](https://help.tableau.com/current/server-linux/en-us/sug_plan.htm)

## 1. Enable REST API and Metadata API

Tableau's [REST API](https://help.tableau.com/current/api/rest_api/en-us/REST/rest_api_requ.htm) and [Metadata API](https://help.tableau.com/current/api/metadata_api/en-us/index.html) **must** be enabled in order to see Tableau metadata in Select Star.

The REST API is permanently enabled for Tableau Cloud, and is enabled by default on all supported versions of Tableau Server.

To enable the Metadata API, a server admin must enable the Metadata API on Tableau Server using the `tsm maintenance metadata-services enable` command through the Tableau Services Manager (TSM) command line interface (CLI). For more information, see [Tableau Documentation](https://help.tableau.com/current/api/metadata_api/en-us/docs/meta_api_start.html#enable-the-tableau-metadata-api-for-tableau-server).

For both Tableau Cloud and Tableau Server, go to **Settings > Automatic Access to Metadata about Databases and Tables**, and make sure the checkbox is checked.

<figure><img src="/files/D4Sb0W88Ooj4PwEfIi91" alt=""><figcaption><p>Automatic Access to Metadata about Databases and Tables</p></figcaption></figure>

{% hint style="info" %}
**Tableau Cloud** can take some time to populate the Metadata Database if you have a big amount of Workbooks in Tableau Cloud.
{% endhint %}

## 2. Create a Personal Access Token

Navigate to **My Account Settings** on the top right of the screen.

![My Account Settings](/files/TeuqLewpc3hoklGBOsg0)

Scroll down to **Personal Access Tokens**.

Enter a name for your token and click **Create new token.**

{% hint style="info" %}
The user you connect to with Select Star needs a **Site Administrator Explorer** role at minimum in order for Select Start to retrieve all the necessary metadata from the Tableau API.
{% endhint %}

![](/files/1zHnrReBTewCpMdy7mMP)

Copy the **Token Secret** from the popup and store it somewhere safe. You will need it later on during this setup.

## 3. Grant access to Tableau Server database (Tableau Server)

{% hint style="info" %}
**Tableau Server**: This step is only applicable if you have a Tableau Server instance. Jump to Step 4 if you are connecting Tableau Cloud.
{% endhint %}

Tableau Server uses an [External Repository](https://help.tableau.com/current/server/en-us/server_external_repo.htm) to store *data about all user interactions, extract refreshes, and more*. **Select Star needs access to this External Repository in Postgres** to properly generate lineage and popularity for Tableau and its data sources.

To provide access, create a new database user with access to all of the following tables in the `public` schema (table and row level):

* public.hist\_users
* public.hist\_views
* public.historical\_events
* public.system\_users
* public.users
* public.views

{% hint style="warning" %}
Note for advanced Postgres setups with **Row-Level Security (RLS)**

\
If your PostgreSQL instance has Row-Level Security (RLS) enabled on any of the above tables (e.g. users, views), the new user may not see any data — even with SELECT permissions.

```
-- Check if RLS is enabled
SELECT relname, relrowsecurity FROM pg_class WHERE relname IN ('users', 'views');
```

In that case, you’ll need to grant access at the row level by adding a policy like:

```
-- For the users table
CREATE POLICY selectstar_users_access
ON public.users
FOR SELECT
TO selectstar_user
USING (true);

-- For the views table
CREATE POLICY selectstar_views_access
ON public.views
FOR SELECT
TO selectstar_user
USING (true);
```

{% endhint %}

{% hint style="info" %}
If your PostgreSQL External repository is located in a private network (Firewalls, SSH, VPN, etc), we also offer the option to sync this information from Snowflake. All we need is the same information transferred to your Snowflake instance.

Contact <support@getselectstar.com> if you have any questions regarding this step.
{% endhint %}

## 4. Create a Tableau Data Source

1. Go to the Select Star **Settings > Data**
2. Click **+ Add** to create a new Data Source.

Fill out the set up form with the information as outlined below.

* **Display Name:** Give your data source a meaningful name, the default is Tableau.
* **Access Token Name:** The Personal Access Token Name.
* **Access Token Secret:** The Personal Access Token Secret.
* **Base Url:** The main URL for your site, such as `https://10ax.online.tableau.com`
* **Site Id:** The name of your site in tableau. You can extract this from the URL you access in Tableau. Example, ***analytics*** is the name of the site for the URL `https://10ax.online.tableau.com/#/site/analytics/home`.

Click the **Connect** button at the bottom of the modal.

<figure><img src="/files/iySXTJh38qD0fAVATtKY" alt=""><figcaption></figcaption></figure>

## 5. Connect Tableau External Repository (Tableau Server)

To connect Tableau External Repository, use the details of the user, database, and schema that you granted access to in Step 3 for Tableau Server. Select Star will connect to

Introduce the details in the setup form as shown in the example below.

<figure><img src="/files/vCx1xNtH03X0seB0Bky9" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
**IP Whitelisting**\
For any instance located in a private network, you must either

* whitelist the following two IP addresses to connect Select Star: `3.23.108.85`, `3.20.56.105`
* Configure the External Repository exporting your activity data to Snowflake. Reach out to <support@getselectstar.com> for more information about this step.
  {% endhint %}

Your metadata should start loading automatically. Please allow 24-48 hours to completely generate popularity and lineage.

When the sync is complete, you'll be able to explore Tableau in Select Star. See the link below for more information on Tableau in Select Star.


# Microsoft Power BI

Follow these steps to connect your Microsoft Power BI instance to Select Star.

## Before you start

To connect Microsoft Power BI to Select Star, you will need...

* Admin access to Microsoft Entra ID (formerly Azure AD)
* Admin access to Microsoft Fabric (formerly Power BI Admin Portal)

Complete the following steps to enable metadata, lineage, and popularity of your Microsoft Power BI in Select Star.

1. [Create an Azure app](#id-1.-create-an-azure-app)
2. [Create a security group in Microsoft Entra ID](#id-2.-create-a-security-group-in-microsoft-entra-id)
3. [Enable the Power BI service admin settings](#id-3.-enable-the-power-bi-service-admin-settings)
4. [Add Azure app to your workspace](#id-4.-add-azure-app-to-your-workspace)
5. [Connect Power BI API to Select Star](#id-5.-connect-power-bi-api-to-select-star)

## 1. Create an Azure app

1\. Open the [Azure Portal](https://ms.portal.azure.com/#allservices) and sign in.

2\. Search for **App registrations**, and select it.

3\. Click **New registration**.

4\. Fill in the required information:

* Name: type "Select Star"
* Supported account types: leave the default value
  * "Accounts in this organizational directory only (xxxxxxxxx only - Single tenant)"
* Redirect URI - leave empty

5\. Click **Register**.

![Screenshot shows the filled registration form.](/files/cdqXBoGQQAfcUjMvp7l5)

6. From the Overview page, copy the **Application (client) ID** and **Directory (tenant) ID**, and securely store them for next steps.

![Screenshot shows where to obtain the Client ID, and Tenant ID](/files/JVdd5SmWzIgxZc4P92es)

7\. Click the **Certificates & secrets** from the left menu.

8\. Under Client secrets, click **+ New client secret**.

![Screenshot shows the new client secret button in the Certificates and secrets tab.](/files/r7fKruIt6C7rkO36lfsS)

In the Add a client secret window, enter a description, select an expiry time, and click **Add**.

* *Sample description: Secret used to connect Select Star to Microsoft Power BI*

Copy the client secret **Value** and securely store it for the next steps.

![Screenshot shows where to obtain the Client secret value in the Certificates and secrets tab](/files/CdktzM4ACLgt6gFtIvDp)

## 2. Create a security group in Microsoft Entra ID

1\. Open the [Azure Portal](https://ms.portal.azure.com/#allservices) and sign in.

2\. Search for **Microsoft Entra ID**, and select it.

3\. Click the **Groups**, under Manage section.

![Screenshot shows the Groups tab for an app in the Azure Active Directory section.](/files/8E38aOHKmF9bjUNPVP3f)

4\. Click **New group**.

5\. Fill in the required information:

* Group type - select "Security"
* Name - type "Power BI - API Access"
* Group description - enter any description or leave empty
  * *Sample description: Security group to grant API access*

6\. Click "No members selected" to open a drawer. Search for **Select Star** user and select it. Click the **Select** button to confirm.

7\. Click **Create**.

![Screenshot shows the filled in form, and selected memeber for security group.](/files/DI6d7OH7iZFFizC3BMf3)

By the end of these steps, you have registered an application with Microsoft Entra ID and created a Security Group with the appropriate member.

## 3. Enable the Power BI service admin settings

1\. Open [Power BI admin portal](https://app.powerbi.com/admin-portal/) and sign in.

2\. Click **Tenant Settings** under the Admin Portal.

* You must have admin access to Microsoft Fabric to configure these settings

3\. Under **Developer settings**:

* Expand **Service principals can call Fabric public APIs**
  * Set this to Enabled.
  * Add your security group you created in [Step 2](#id-2.-create-a-security-group-in-microsoft-entra-id), under **Specific security groups.**
  * Click **Apply**.

4\. Repeat the process for the subsections under the **Admin API settings** section.

* Open the section, Set **Enabled,** and add the security group you created in [Step 2](#id-2.-create-a-security-group-in-microsoft-entra-id). Click **Apply**.

You must complete the steps for the following sections under **Admin API settings**:

* Service principals can access read-only admin APIs
* Enhance admin APIs responses with detailed metadata
* Enhance admin APIs responses with DAX and mashup expressions

<figure><img src="/files/B6DvBPx9gbPoFEDtcvs8" alt=""><figcaption><p>The screenshot shows the highleted sections, needed to enabled for Microsoft Power BI Integration</p></figcaption></figure>

## 4. Add Azure app to your workspace

1\. Open [Power BI](https://app.powerbi.com/) and sign in.

2\. Search for the workspace you want to enable access for, and from the three-button menu, select **Workspace access**.

![Screenshot shows the "Workspace access" button in workspace list.](/files/SHTTsvO1inxcPBlLPRNP)

3\. Click **+ Add people or groups**

4\. Search for the app you created in [step 1](#id-1.-create-an-azure-app), i.e, **Select Star**, and select it. Set the permissions to **Contributor**.

5\. Click **Add**, and close the drawer.

6\. Repeat the above steps for all workspaces you want to be added to Select Star.

{% hint style="warning" %}
**Important!** If you have any Power BI reports using Semantic Models from other workspaces, please make sure to add the Azure App you created in [**Step 1**](#id-1.-create-an-azure-app) as a **Contributor** to all those workspaces.\
This is required to ingest the reports' metadata and generate the column-level lineage.
{% endhint %}

## 5. Connect Power BI API to Select Star

1\. Go to the Select Star **Settings**. Click **Data** in the sidebar, then **+ Add** to create a new Data Source.

![Screenshot shows the "Add Data source" tab for PowerBI in Select Star.](/files/n0dpRuVVZWNC5WY5BFqW)

2\. Fill in the required information:

* **Client ID:** Application (client) ID of Azure App from [step 1.6](#id-1.-create-an-azure-app) above.
* **Client Secret:** Client secret value of Azure App from [step 1.8](#id-1.-create-an-azure-app) above.
* **Tenant ID:** Directory (tenant) ID of Azure App from [step 1.6](#id-1.-create-an-azure-app) above.


# Looker

Follow these steps to connect your Looker instance to Select Star.

## Before you start

To connect Looker to Select Star, you will need...

* Admin access to your Looker account.
* Admin access to your LookML repo(s) in GitHub.

Complete all of the following steps to see Looker metadata, lineage, and popularity in Select Star.

{% hint style="info" %}
If you want to **create an API3 Key** on a **Looker account with Admin access**, start at [step 3](#3.-create-a-new-api3-key).
{% endhint %}

1. [Create a Select Star Permission Set and Role in Looker](#1.-create-a-select-star-permission-set-and-role-in-looker)
2. [Create a Select Star user in Looker](#2.-create-a-select-star-user-in-looker)
3. [Create a new API3 Key](#3.-create-a-new-api3-key)
4. [Grant view access to Shared Folder](#4.-grant-view-access-to-shared-folder)
5. [Connect Looker API in Select Star](#connect-looker-to-select-star)
6. [Connect LookML repo to Select Star](#connect-lookml-repo-to-select-star)

## 1. Create a Select Star Permission Set and Role in Looker

Create a new Permission Set called `Select Star` with the following permissions:

![](/files/WNyU8y8FHYfcDh1uH3Ey)

![](/files/DJfgVNJXX0gVUz1dRREY)

![](/files/vCjULCya633AML4Ep6m0)

Create a new role called `Select Star Role` . Choose the `Select Star` Permission Set you just created and the `All` Model Set as shown below.

![](/files/SIvRpikJRxS8rf4f2Oua)

When created, the Role Permissions will look like the following:

![](/files/gC7aD1ylAbPA2w0DFvL8)

## 2. Create a Select Star user in Looker

Now the role is created, you can assign the `Select Star Role` to a Looker User.

You can either assign the role to an existing Looker user, or create a new user by sending an invite to <selectstar@getselectstar.com> with the `Select Star Role`.

{% hint style="info" %}
You can also create an API3 key in your own Looker account, as long as your permissions include everything from step 1.
{% endhint %}

## 3. Create a new API3 Key

From the main Looker page, Click **Admin**, then **Users**.

![](/files/53RJuLdzmvIGZiElAATM)

Find your Select Star user and click the **Edit** button. Under API3 Keys, click the **Edit Keys** button.

![](/files/Ph1HGAfwbnofy4njHtFA)

Create a **New API3 Key**.

![](/files/Ej85CdhNVcuyW8Cm4yAR)

You will see a new API3 Key with a **Client ID** and **Client Secret**.

## 4. Grant view access to Shared Folder

You may need to grant view access to the user you created in [step 1](#1-create-select-star-user-in-looker) on the Shared Folder in order to see your dashboards in Select Star.

Navigate to the Shared folder in Looker.

![](/files/-MkiCyh6NCHn2hslOQkb)

Find the gear icon in the upper right part of the screen and click **Manage Access.**

![](/files/-MkiD60NlrKgHCkMCt9l)

Grant **View** access to the Select Star user you created.

![](/files/-MkiDGrnv3ZMBS4R0fuE)

## 5. Connect Looker API to Select Star

Go to the Select Star **Settings**. Click **Data** in the sidebar, then **+ Add** to create a new Data Source.

![](/files/IP5pzA0NjiBfKRqyS8wu)

Choose **Looker** in the Source Type dropdown and provide the following information:

![](/files/Ps3g7BdYvNvjyU3zsqtV)

* **Display Name:** This value is `Looker` by default, but you can override it if desired.
* **API Client ID:** The API3 Client ID from [Step 4](#4.-grant-view-access-to-shared-folder) above.
* **API Client Secret:** The API3 Client Secret from [Step 4](#4.-grant-view-access-to-shared-folder) above.
* **Host URL:** The full URL of your Looker instance.

Click **Connect** to proceed.

## **6. Connect LookML repo to Select Star**

In order for Select Star to connect the data warehouse and Looker data models, access to the GitHub repo for the LookML code is required. If you do not need to see lineage, you can click **Next** to skip this step.

Select Star will automatically detect LookML projects connected to your Looker instance.

{% hint style="info" %}
If you do not see any LookML projects, please check the Model Set in [Step 1](#1.-create-a-select-star-permission-set-and-role-in-looker). Select Star detects LookML projects based on access to Looker models using the API.
{% endhint %}

{% hint style="warning" %}
You must use a SSH connection when integrating Looker with Git. This is required to set up deploy keys on most major Git Repositories including Github and Gitlab.

Check the *Connecting to Git using SSH* section in the Looker docs.

<https://docs.looker.com/data-modeling/getting-started/setting-up-git-connection>
{% endhint %}

![](/files/wERZq0ydb9EZl3vkfGzf)

For each project you want to connect, click **Copy Key**.

Open your browser in a new window or tab, and go to your LookML repo in GitHub.

Click on **Settings >** **Deploy keys** on the sidebar.

Click the **Add deploy key** button in the corner.

<figure><img src="/files/u3p9rtRwCe1KQZoi1C5C" alt=""><figcaption></figcaption></figure>

Paste in the entire key copied from Select Star into the **Key** text box.

Set the title to a descriptive name, such as `Select Star Ingestion Key`.

**You do not need to allow write access.**

Click **Add key.**

![](/files/WFmMgAH4u6HpbuwkWY57)

Return to Select Star and click the **Next** button at the bottom of the modal.

Your metadata should start loading automatically. Please allow 24-48 hours to completely generate popularity and lineage.

![](/files/z09nRr1FbyLCs8gyqDK6)

When the sync is complete, you'll be able to explore Looker in Select Star. See the link below for more information on Looker in Select Star.

{% content-ref url="/pages/-MgSXq1Niws-e6rtFJ\_P" %}
[Getting Started: Looker](/learning-data/getting-started-looker)
{% endcontent-ref %}


# Metabase

Follow these steps to connect your Metabase to Select Star.

1. [Create an API Key](#id-1.-create-an-api-key)
2. [Connect Metabase to Select Star](#id-2.-connect-metabase-to-select-star)
3. [Adjusting connections](#id-4.-adjusting-connections)

## 1. Create an API Key

1. Click the gear icon (⚙️) in the top right, then select **Admin settings**.
2. In the left-hand menu, click on **Settings**, then **Authentication**.
3. Scroll down to the **API Keys** section and click the **Manage** button.
4. Click the **Create API Key** button.
5. Create a new API Key:
   * **Key name:** Give it a descriptive name.
   * **Select a Group:** Select **Administrator**.

![](/files/Xk1iU0ysLcCMRJ7P7X9T)

Metabase will generate and display the API key. Store this key securely, you won't have access to it again later, and if you lose it, you'll need to regenerate it (which invalidates the old one).

{% hint style="info" %}
We need **Admin** permissions to fetch all users from Metabase and the complete lineage. With the limited set of permissions it is not possible.
{% endhint %}

## 2. Connect Metabase to Select Star

Next, you will add Metabase as a data source. Go to the Select Star **Settings.** Click **Data** in the sidebar, then **+ Add** to create a new Data Source.

![](/files/IP5pzA0NjiBfKRqyS8wu)

Choose Metabase in the Source Type dropdown and provide the following information:

![Create Metabase Datasource in SelectStar](/files/sYbLXTRsR5PWvkBwxkr9)

* **Display Name:** The name you want to give to your new data source.
* **Url:** The URL of your Metabase instance.
* **API Key:** The API Key generated on step one.

Click **Save** to proceed.

### Choose Your Metabase Deployment Type

Specify whether your Metabase instance is cloud-based or on-prem.

![Screenshot shows choice of cloud vs on-prem](/files/VxtpOAe7kDBtKvfkSb56)

#### Cloud

If you select **Cloud**, simply click **Save**. This option is also suitable for self-hosted Metabase Pro and Enterprise plans where Metabase manages its own database.

Upon saving, your Metabase metadata will begin loading automatically into Select Star. Please allow 24-48 hours for complete indexing, popularity score generation, and lineage mapping. Once synchronization is complete, you can start exploring your Metabase data within Select Star.

#### On-Prem

1. Create a dedicated PostgreSQL database user that Select Star will use to access your Metabase's internal database.

   ```sql
   create user selectstar;
   alter user selectstar password '<password>';
   grant connect on database metabase to selectstar;
   grant select on table query_execution to selectstar;
   ```
2. Provide details of the PostgreSQL database where your Metabase instance stores its internal information. After entering the database credentials, click **Save**.

{% hint style="info" %}
**IP Whitelisting:** If your instance is protected by a firewall, you must whitelist the following Select Star IP addresses to enable connection: `3.23.108.85, 3.20.56.105`
{% endhint %}

![Screenshot shows on-prem database credentials](/files/r3epfttkEqWg3aaUnAY7)

Your metadata will begin loading automatically. Please allow 24-48 hours for complete indexing, popularity score generation, and lineage mapping. Once the synchronization is complete, you can explore your Metabase data within Select Star.

{% hint style="info" %}
If you require integration with a different database type, please contact our sales team.
{% endhint %}

## 3. Adjusting connections

Select Star automatically attempts to detect connection types for your Metabase connections. However, if you are connecting to different instances or data sources, you may need to adjust these mappings manually.

1. Go to **Settings**
2. Click on **Data**
3. Click on your **Metabase** data source
4. Switch to the **Connections** tab
5. Make sure the checkbox for the connection is checked
6. Make sure the **Select Star Datasource** is correctly selected

![Screenshot shows on-prem database credentials](/files/fOY2ZZaUdrEccZ5H2jz2)


# Fivetran (beta)

Follow these steps to connect your Fivetran instance to Select Star.

{% hint style="warning" %}
Important: The Fivetran integration requires you to have **Fivetran Enterprise** or above.
{% endhint %}

Select Star uses [Fivetran Platform Connector](https://fivetran.com/docs/logs/fivetran-platform) to fetch metadata and lineage information of your Fivetran instance.

{% hint style="info" %}
Select Star currently does not catalog the metadata of the Fivetran connectors, and only generates lineage information for the assets synced by Fivetran.
{% endhint %}

## Before you start

To connect Fivetran to Select Star, you will need...

* Admin access to your Destination/Data Warehouse instance. (to configure the permissions)
* Admin access to your Fivetran account. i.e., user account with `Account Administrator` role. (to configure the Fivetran Platform Connector)

Complete all the following steps to see lineage from Fivetran in Select Star.

1. [Configure Fivetran Platform Connector](#id-1.-configure-fivetran-platform-connector)
2. [Connect Fivetran to Select Star](#id-2.-connect-fivetran-to-select-star)

## 1. Configure Fivetran Platform Connector

Follow Fivetran's step-by-step [setup guide](https://fivetran.com/docs/logs/fivetran-platform/setup-guide) to manually set up your Fivetran Platform connection account-wide. Skip this step if you have already set up the account-wide Fivetran Platform Connector.

Select Star currently supports the following destinations to connect with Fivetran:

* Snowflake
* Google BigQuery
* PostgreSQL

If you are using a different destination, please reach out to us, and our dedicated support team will help you with the integration.

## 2. Connect Fivetran to Select Star

Next, you will add Fivetran as a data sources, and specify the destination that the Fivetran Platform Connector has delivered your logs to.

Go to the Select Star **Settings**. Click **Data** in the sidebar, then **+ Add** to create a new Data Source.

![](/files/IP5pzA0NjiBfKRqyS8wu)

Choose **Fivetran** in the Source Type dropdown and provide the following information:

![](/files/och4f9Vm8lcY2lcDxMhd)

* **Display Name** - The name you want to give to your new data source.

Click **Save** to proceed.

#### Configure Destination Type

Select the destination where you are syncing your Fivetran Platform Connector logs.

* **Destination Type** - The destination type where Fivetran is syncing the data.

![](/files/lzzCNcpMKNKuMIZey2Hd)

Please follow the instructions below and fill in the required details based on the destination type you selected.

* [Snowflake](/integrations/snowflake)
* [Google BigQuery](/integrations/bigquery)
* [PostgreSQL](/integrations/postgres)

{% hint style="warning" %}
IMPORTANT! Please ensure that the user configured in the destination for Select Star has the SELECT permissions on all the tables in the schema used by Fivetran Platform Connector.

These tables contain the metadata and lineage information that Select Star uses to generate the lineage. For more information, refer to the [Fivetran Platform Connector](https://fivetran.com/docs/logs/fivetran-platform) documentation.
{% endhint %}

Click **Save** to proceed.

{% hint style="info" %}
Select Star currently matches schema names in the destination with those in Fivetran to identify the correct database and generate lineage.

If you believe the lineage is not being generated correctly, please reach out to our support team.
{% endhint %}

Your metadata should start loading automatically. Please allow 24-48 hours to completely generate the lineage.


# Mode

The Select Star integration allows users to view, search and understand their Mode instance all in one convenient place. Discover the benefits and features of the Mode Select Star integration in this

## Before you start

To connect Mode to Select Star, you will need...

* Admin access to your Mode account.
* A paid Mode plan (required for API token generation).
* If you want to see Mode reports in personal spaces and calculate popularity based on the last 90 days of user activity, you will need credentials for your Mode Discovery Database. Learn more about getting access to Mode Discovery Database [here](https://mode.com/developer/discovery-database/introduction/).

Complete all of the following steps to see Mode metadata, lineage, and popularity in Select Star.

1. [Create an API Access Token](#1.-create-an-api-access-token)
2. [Connect Mode API to Select Star](#2.-connect-mode-to-select-star)
3. [Connect Mode Discovery Database to Select Star](#3.-connect-mode-discovery-database)

## 1. Create an API Access Token

To enable integration with Mode, you can use either Workspace API tokens or Member API tokens. However, we strongly recommend using Workspace API tokens as they provide comprehensive access and help avoid permission issues when organizational structures change.

### Workspace API Tokens

Workspace API tokens are designed for programmatic management of your Mode workspace. They mimic Admin access to the Workspace and are ideal for tasks such as archiving inactive reports or managing schedules. Only Admins can create and manage these tokens.

To generate a Workspace API token:

1. Navigate to **Workspace Settings** > **Features** > **API Keys** > **Workspace keys**.
2. Click "Create API Key" and provide a display name.
3. Save "Key ID" and "Secret" values, as you will need them in [step 2](#2.-connect-mode-to-select-star).

### Member API Tokens

Member API tokens are intended for individual use cases, such as updating specific reports or managing collections. These tokens mimic the individual user's access to resources in the Workspace.

To generate a Member API token:

1. Admins must enable Member API tokens in **Workspace Settings** > **Features** > **API Keys** > **Member keys**.
2. Users can then create their own tokens in **Workspace Settings** > **Personal** > **My API Keys**.
3. Save "Key ID" and "Secret" values, as you will need them in [step 2](#2.-connect-mode-to-select-star).

### Important Notes

* API tokens expire after 90 days by default.\
  – Using a Member API token allows you to restrict access to specific collections within Select Star, but it introduces additional complexity to the permissions model.
* Ensure you securely store both the Token (Key ID) and Secret when creating any API token.
* For more details, refer to [Mode's API token documentation](https://mode.com/help/articles/api-tokens).

## 2. Connect Mode API to Select Star

Go to the Select Star **Settings**. Click **Data** in the sidebar, then **+ Add** to create a new Data Source.

![](/files/IP5pzA0NjiBfKRqyS8wu)

Choose **Mode** in the Source Type dropdown and provide the following information:

![](/files/KPYqkjAz81DZe68sGvDo)

* **Display Name:** This value is `Mode` by default, but you can override it if desired.
* **API access token:** The `Key ID` value from the API token created in Mode.
* **Token secret:** The `Secret` value from the API token created in Mode.
* **Organization:** The name of your Mode organization. It'll be the part of the url after `app.mode.com/home/`
* **DB Connection:** The database you are running Mode on top of, such as Snowflake or BigQuery.
* **DB Username:** The name of the service account user that runs mode queries in your DB Connection. It is not a Select Star service account. This is used to identify which queries Mode runs so that Select Star can generate lineage between the database and Mode.

Click the **Connect** button.

{% hint style="success" %}
You must ensure DB Connection and DB Username are set so we can calculate column-level lineage from your DWH to Mode dashboards. You can also verify wether or not the connection has been set by going to [**Settings > Data**](https://app.selectstar.com/admin/data) **> Mode > Connections**.
{% endhint %}

## 3. Connect Mode Discovery Database

{% hint style="info" %}
This step is optional, but **highly recommended** if you want to see Mode reports in personal spaces and to calculate popularity scores based on the last 90 days of user activity.

To learn how to get the Snowflake credentials for Mode Discovery Database, see Mode's instructions [here](https://mode.com/developer/discovery-database/introduction/).
{% endhint %}

To skip this step, click the **Next** button.

![](/files/oqlFAa6l1W9ItsMoavoT)

You'll need the following information:

* **Account:** Your Mode Discovery Database account name.
* **Username:** Your Mode Discovery Database username.
* **Password:** Your Mode Discovery Database password.
* **Role:** By default, this value is `PUBLIC`.
* **Warehouse:** The name of your Mode Discovery Database. By default, this value is `WAREHOUSE_XS_MODE`.
* **Database:** By default, this value is `MODE`.
* **Schema:** By default, this value is `ORGANIZATION_USAGE`.

When you have entered the information, click **Next**.

Your metadata should start loading automatically. Please allow 24-48 hours to completely generate popularity and lineage.

When the sync is complete, you'll be able to explore Mode in Select Star. See the link below for more information on Mode in Select Star.

{% content-ref url="/pages/-MgSXtXm-BTwgdI5WQBd" %}
[Getting Started: Mode](/learning-data/getting-started-mode)
{% endcontent-ref %}


# Sigma Computing

The Select Star integration allows users to view, search and understand their Sigma instance all in one convenient place.

## Before you start

To connect Sigma to Select Star, you will need...

* Admin access to your Sigma account.

Complete the following steps to enable metadata, lineage, and popularity of your Sigma workbooks in Select Star.

1. [Identify your Sigma API request URL](#id-1.-identify-your-sigma-api-request-url)
2. [Generate API Client ID and Secret](#id-2.-generate-api-client-id-and-secret)
3. [Connect Sigma API to Select Star](#id-3.-connect-sigma-api-to-select-star)

## 1. Identify your Sigma API request URL

The URL Select Star uses to send API requests depends on the [cloud your Sigma organization is hosted on](https://help.sigmacomputing.com/docs/region-warehouse-and-feature-support).

View your **Base URL** in the **Administration** section of Sigma. Go to **Administration** > **Developer Access** > **API base URL**. Keep this information for the next steps.

<figure><img src="/files/EeLVRDxqNd6dQmp2lz9m" alt=""><figcaption></figcaption></figure>

## 2. Generate API client ID and secret

To generate API client credentials, you must be assigned the **Admin** account type.

1. From **Sigma Home**, open **Administration**, or click your user avatar to open the user menu and select **Administration**.
2. In the side panel, select **Developer Access**.
3. Click **Create new** to set up new credentials.
   1. The **Create client credentials** modal opens.
4. For **Select scopes**, select the REST API checkbox to enable the use of these credentials for the API.
5. For **Name**, enter a unique name to identify the credentials. For example, "**Select Star**".
6. \[optional] For **Description**, enter a description of the purpose of the credentials. For example, "Token used to connect Select Star integration"
7. For **Owner**, set yourself as the owner, OR search for and select a member of your organisation with whom to associate the credentials.
   1. Note: The API secret uses the account type permissions assigned to this user.
   2. **Select Star recommends selecting an Admin as the Owner in order to fetch all workbooks available in your Sigma instance.**
8. Click **Create** to generate the credentials.

![](/files/811FgnUqSuMQBmfUbAFT)

9. Copy the client ID and secret, and securely store them for the next steps.

![](/files/xE3eoFp4A1KMYuJrGvtl)

## 3. Connect Sigma API to Select Star

Go to the Select Star **Settings**. Click **Data** in the sidebar, then **+ Add** to create a new Data Source.

Choose **Sigma** in the Source Type dropdown and provide the following information:

![](/files/cdKpVeq1KUmU9dKvtMBZ)

* **Display Name:** This value is `Sigma` by default, but you can override it if desired.
* **API Client ID:** The API Client ID from [step 2](#id-2.-generate-api-client-id-and-secret) above.
* **API Client Secret:** The API Secret Token from [step 2](#id-2.-generate-api-client-id-and-secret) above.
* **API Server:** This is based on your Sigma cloud configuration. You need to use the API request URL from [step 1](#id-1.-identify-your-sigma-api-request-url) above.
* **DB Connection:** This is based on the database connection used in your Sigma instance.
* **DB Username:** This is the database user configured in your Sigma instance.

{% hint style="info" %}
In order to show the cross-platform lineage, the DB Connection & DB Username need to match with your Sigma connection details.
{% endhint %}

![Sigma Administration page](/files/0qNAdzDxVqPZa76Pb6pq)

{% content-ref url="/pages/hPMwCThQDwCcn8sD4m2R" %}
[Getting Started: Sigma](/learning-data/getting-started-sigma)
{% endcontent-ref %}


# Sisense / Periscope (beta)

Follow these steps to connect your Sisense/Periscope instance to Select Star.

## Before you start

To connect Periscope to Select Star, you will need...

* Active Periscope Account email and password credentials or Google Account credentials, for Google SSO-enabled accounts
* To calculate lineage and popularity based on the last 90 days of users' activity, you'll need to connect Select Star to the Data Warehouse that your Periscope instance is using to run queries.
* To calculate popularity and top users, Select Star requires a mapping of periscope users to their email addresses, you'll need to prepare a CSV file with a `user name: email` mapping.

Complete all of the following steps to see Periscope metadata, lineage, and popularity in Select Star.

1. [Create an Account Authentication Token](#1.-create-an-account-authentication-token) (optional, only if you're using Google Single Sign-On)
2. [Connect Periscope to Select Star](#2.-connect-periscope-to-select-star)
3. [Import User Catalog](#3.-import-user-catalog)

## 1. Create an Account Authentication Token

**Optional, only for Google SSO-enabled accounts**

Prepare your Google Account credentials and contact Select Star support to help you with the setup process.

After completing the procedure, you should receive a file named `periscope-credentials.json`. Please save it, as you will need this file in [step 2](#2.-connect-periscope-to-select-star).

## 2. Connect Periscope to Select Star

Go to the Select Star `Settings`. Click `Data` in the sidebar, then `+ Add` to create a new Data Source.

![](/files/jufHHtK9m7TEJCX480a7)

Choose `Sisense/Periscope` in the Source Type dropdown and provide the following information:

For email / password authentication please leave the default `Email and Password Credentials` checked. If your account uses Google SSO to authenticate, check `Cookies json`.

### Email and Password

![](/files/5JkTFxKRQO76qqVZTMTE)

* **Display Name:** This value is `Sisense / Periscope` by default, but you can override it if desired.
* **Email:** Your Periscope Account email
* **Password:** Your Periscope Account password
* **DB Connection:** The database you are running Periscope on top of.
* **DB Username:** The name of the service account user that runs Periscope queries in your DB Connection. It is not a Select Star service account. This is used to identify which queries Periscope runs so that Select Star can generate lineage between the database and Periscope.

### Google SSO Authentication

![](/files/aR9wi8rjpIpyCg3c4MJ6)

* **Display Name:** This value is `Sisense / Periscope` by default, but you can override it if desired.
* **Google Auth credentials JSON file:** The `periscope-credentials.json` file created in [step 1](#1.-create-an-account-authentication-token).
* **DB Connection:** The database you are running Periscope on top of.
* **DB Username:** The name of the service account user that runs Periscope queries in your DB Connection. It is not a Select Star service account. This is used to identify which queries Periscope runs so that Select Star can generate lineage between the database and Periscope.

Click the `Connect` button.

Your metadata should start loading automatically. Please allow 24-48 hours to completely generate popularity and lineage.

When the sync is complete, you'll be able to explore Periscope in Select Star.

## 3. Import User Catalog

{% hint style="info" %}
Although this step is optional, you won't see popularity / top users unless the catalog is imported.
{% endhint %}

To complete the import, please prepare a CSV file with at least the following columns:

* First Name
* Last Name
* Email

Alternatively:

* Full Name
* Email

and then contact Select Star support, <support@getselectstar.com> to do the import for you.


# Looker Studio (beta)

The Select Star integration allows users to view, search and understand their Looker instance all in one convenient place. Discover the benefits and features of the Looker Select Star integration in t

## Before you start

To connect Google Data Studio to Select Star, you will need...

* Access to Google Workspace admin console
* Access to Google Cloud project for BigQuery exports

1. [Extract session](#1.-extract-session)
2. [Setup Google Workspace exports](#2.-setup-google-workspace-exports)
3. [Setup GCP service account for accessing exports](#3.-setup-gcp-service-account-for-accessing-exports)
4. [Connect Google Data Studio API to Select Star](#4.-connect-google-data-studio-api-to-select-star)

## 1. Extract session

{% hint style="warning" %}
For the best outcome, we strongly recommend you follow the steps below using an new incognito window to sign into Looker Studio.
{% endhint %}

Google Data Studio does not currently provide APIs that are necessary for the functioning of Select Star. For this reason, Select Star is forced to intercept the user account.

To provide access to Google Data Studio, we recommend to [create a dedicated user account](https://support.google.com/a/answer/33310?hl=en) for Select Star that serves as a service account. However, it is not necessary and not required, especially during the trial and evaluation stage.

1\. Open "Incognito Mode" in your web browser. Perform this step even if you are currently logged in to the desired user account to obtain its cookies correctly.

2\. Sign to a Google user account with access to Google Data Studio.

3\. Open [Google Data Studio dashboard](https://datastudio.google.com/) in Google Chrome.

4\. **Open Developer console in Google**. To open the developer console in Google Chrome, open the *Chrome Menu* in the upper-right-hand corner of the browser window and select *More Tools* > *Developer Tools*. You can also use `Option` + `⌘` + `J` (on macOS), or `Shift` + `CTRL` + `J` (on Windows/Linux).

5\. Select tab "Network" in the Developer console. Refresh the browser tab to fill the "Network" tab with data. Search for request to "datastudio.google.com".

6\. Find a successful request that is sent to `lookerstudio.google.com`, then open context menu and select **Copy as cURL**

7\. Keep this value safe as you will need it to create integrations in final steps.

<figure><img src="/files/pbm9YqLkdmkkCdRJilEx" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
If the Domain column (or any other) is not displayed in your Network tab, you can enable it by right clicking the header of the request table, and making sure the corresponding column is checked.
{% endhint %}

The curl request should look like the one below

```sh
curl 'https://lookerstudio.google.com/batchedDataV2?appVersion=20231211_0700' \
  -H 'authority: lookerstudio.google.com' \
  -H 'accept: application/json, text/plain, */*' \
  -H 'accept-language: en-US,en;q=0.9' \
  -H 'cache-control: no-cache' \
  -H 'content-type: application/json' \
  -H 'cookie: [A VERY LONG STRING]
  .... MORE HEADERS ...
  --data-raw '[A VERY LONG STRING]' \
  --compressed
```

## 2. Setup Google Workspace exports

To set up a BigQuery Export configuration, you first need to set up a BigQuery project in the Google Cloud console. To analyze activity, including calculating the popularity of individual reports, it is necessary to provide Select Star access to audit logs and usage reports for Google Data Studio by exporting them to Google BigQuery.

To set up a BigQuery Export configuration, follow Google Workspace documentation about "[Set up service log exports to BigQuery](https://support.google.com/a/answer/9079365)".

Note that the export only contains logs from its setup until it is disabled. Exporting them to BigQuery may take Google up to 48 hours, and then Select Star should ingest them within 24 hours.

## 3. Setup GCP service account for accessing exports

Next create a service account with the permissions to read the project with exports:

* BigQuery Data Viewer
* BigQuery Job User

This service account will be later used for reading the activity logs.

## 4. Connect Google Data Studio API to Select Star

1\. Go to the Select Star **Settings**. Click **Data** in the sidebar, then **+ Add** to create a new Data Source.

<figure><img src="/files/QYjYIFZsHeM9rQ8PKqIu" alt=""><figcaption></figcaption></figure>

2\. Fill form in the required information:

* **CURL command with cookies**: CURL with cookies that you extracted in [step 1.5](#1.-extract-session) of the instructions.

3\. Proceed to next step and select **Skip Activity Source** if you don't want to use popularity feature in Select Star.

<figure><img src="/files/kKLow0OlBpE5iu3WbQXk" alt=""><figcaption></figcaption></figure>

4\. Proceed to next step and select **Add Activity Source** dropdown and fill **Activity Log Source** together with supplying BQ service account json with credentials that you created in [step 3](#3.-setup-gcp-service-account-for-accessing-exports).

<figure><img src="/files/26YEAzQoSsjhErZpxId7" alt=""><figcaption></figcaption></figure>

## 4. Chart names

Pages in Looker Studio can contain one or multiple charts. Looker Studio offers a canvas on which to place those charts, but there is no way to assign a text field to a given chart.

Select Star will figure out automatically the chart titles by looking at text fields that are positioned immediately before a given chart.

To ensure your Looker Studio connection to Select Star properly ingests the names for all the Charts in your Pages, please make sure that you add a text box with the title of the chart above it.

Find an example in the images below of the layout that Select Star expects.

<figure><img src="/files/Y65xaj4ve87mj0wAIoZm" alt=""><figcaption><p>Page view in Select Star</p></figcaption></figure>

<figure><img src="/files/sHtmI5lAuCJCmIs0qGmJ" alt=""><figcaption><p>Page view in Looker Studio</p></figcaption></figure>

## 5. Refreshing data

Looker Studio has limitations that prevent automatic data refreshing. But we have ways to work around that, if you need automatic refresh, please get in touch with [customer support](mailto:support@getselectstar.com). If you want to learn more about this integration we can show quickly you around, just shoot us an [email](mailto:sales@getselectstar.com).


# ThoughtSpot

Follow these steps to connect your ThoughtSpot instance to Select Star.

## Before you start

To connect ThoughtSpot to Select Star, you will need...

1. [Generate a Secret Key](#1.-generate-a-secret-key)
2. [Connect ThoughtSpot to Select Star](#2.-connect-thoughtspot-to-select-star)

## **1. Generate a Secret Key**

Generate a Secret Key in the Develop tab,

1. As an account admin, log in to the ThoughtSpot.
2. Click on thr **Develop tab**.

<figure><img src="/files/hZV3iI0npLY6klNn0Cwp" alt=""><figcaption></figcaption></figure>

3. Under the **Security settings** section, click Edit, to enable `Trusted authentication`.

<figure><img src="/files/0AM9IUycv5rdLId8To36" alt=""><figcaption></figcaption></figure>

4. Toggle the switch to enable the feature and **Save Changes**.

<figure><img src="/files/T88GK4iJyiLhPgz4RqrK" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
Enabling the Trusted authentication, will trigger a service interruption.
{% endhint %}

<figure><img src="/files/AHwIlsJxDQBHguNnAQ9y" alt=""><figcaption></figcaption></figure>

5. After the changes are saved, you should be able to see and copy the **Secret Key**.

<figure><img src="/files/cgcxShL2GUt62iXx0NSK" alt=""><figcaption></figcaption></figure>

## 2. Connect ThoughtSpot to Select Star

1. Go to the Select Star **Settings**.
2. Click **Data** in the sidebar, then **+ Add** to create a new Data Source.

<figure><img src="/files/0xM0kSfRsdJ3cOWgRQos" alt=""><figcaption></figcaption></figure>

3. Choose **ThoughtSpot** in the Source Type dropdown and provide the following information:

**Display Name:** This value is `ThoughtSpot` by default, but you can override it if desired.

**Username:** This is the username of the user who created the Secret Key from Step 1.

**Secret Key:** This is the **Secret Key** from Step 1, which is used to authenticate access to Thoughtspot.

**Base URL:** This is the address of your ThoughtSpot instance. This should include the `https://<instance name>.thoughtspot.cloud`.

<figure><img src="/files/pbv9VEVu5BJXU52dym4w" alt=""><figcaption></figcaption></figure>

4. Once this information is added, click Connect. This will trigger a sync process. You can close this window now.

<figure><img src="/files/7FSPIkCVaTqOfMTe6ZBG" alt=""><figcaption></figcaption></figure>

Your metadata should start loading automatically. Please allow 24-48 hours to completely generate and lineage.

When the sync is complete, you'll be able to explore ThoughtSpot in Select Star.

{% content-ref url="/pages/rn9RRtC3ipOX2Vw7Lpba" %}
[Getting Started: ThoughtSpot](/learning-data/getting-started-thoughtspot)
{% endcontent-ref %}


# QuickSight (beta)

Learn how to connect QuickSight to Select Star with our step-by-step guide. Discover the benefits of integrating these two powerful platforms for better data management and analysis.

To connect QuickSight to Select Star, you will need to...

1. [Connect QuickSight to Select Star](#1.-connect-quicksight-to-select-star)
2. [Create an AWS IAM Role ARN using Cloudformation](#2.-create-an-aws-iam-role-arn-using-cloudformation)
3. [Confirm authorization](#3.-confirm-authorization)

{% hint style="info" %}

* Select Star requires only minimal metadata access to AWS QuickSight
* The granted permissions are defined in CloudFormation template
* IAM permission is defined by resource "CrossAccountRolePolicy" in file [SelectStarQuickSight.json](https://github.com/selectstar/cloudformation-templates/blob/main/quicksight/SelectStarQuickSight.json)
  {% endhint %}

### **1**. Connect QuickSight to Select Star

Select QuickSight from the Add Data Source menu and provide the

**Display Name** - This value is `QuickSight` by default, but you can override it if desired.

**Region** - ID of the AWS region where the cluster was created. For example `us-east-2`,`us-west-1`, `eu-central-1`

![](/files/GHy2ZDYiAG4Vq3O4jGuh)

### 2. **Create** an AWS IAM Role ARN using CloudFormation

Select Star recommends use of AWS CloudFormation to setup integration, which allows you to make necessary changes to the QuickSight environment in an automatic, transparent, safe, and auditable manner.

AWS CloudFormation creates an AWS IAM Role to enable access for Select Star and add it to QuickSight cluster.

The source code of the CloudFormation template along with build scripts and real-time logs of the continuous deployment system is available on public repository on GitHub "[selectstar/cloudformation-templates](https://github.com/selectstar/cloudformation-templates/tree/main/quicksight)" to be freely audited.

You can pass the link to CloudFormation to the infrastructure team to enable the integration to be created.

1\. A simple form will be displayed in Select Star.

![](/files/wFFVY7nHc0riTAA1z0RT)

2\. Select the "Open CloudFormation" button. A new window will open to proceed to the creation of a CloudFormation stack by AWS Management Console. Make sure you are logged into the AWS account in which the QuickSight is hosted.

3\. The **Create Stack** form will be displayed. Fill it in with the required information:

![](/files/6NBsYO9ibHFClCYTciBJ)

4\. Review the correctness of information in the form. Under **Capabilities** check "I acknowledge that AWS CloudFormation might create IAM resources".

![](/files/buylqYRG2xIWuEVaCcnD)

5\. Click **Create stack**.

6\. Wait until the stack changes it status to "<mark style="color:green;">CREATE\_COMPLETE</mark>" from "CREATE\_IN\_PROGRESS" in tab "**Stack Info**". The operation should take up to 5 minutes. You need to refresh the tab to see the progress.

![](/files/bYb7KgE0ksgrN0yQbb1I)

7\. After completing stack creation, the `Role ARN` is available from the "**Outputs**". Copy and save the `RoleArn` for later use.

![Outputs tab](/files/p38zgGz1GpKK8Z7N0vkV)

### 3. Confirm authorization

1\. Return to Select Star. You should see a form that allows you to provide "Role ARN". Pay attention to its value, it should not contain any whitespace characters (e.g. tab at the end of the line). Fill out the form with the required information:

* **Role ARN:** Identifier of AWS IAM Role to use by Select Star. You'll see this after completing [step 2.7](#2.-create-an-aws-iam-role-arn-using-cloudformation) of the instructions.

![Enter Role ARN from Step 2.7](/files/m8Nje8hjp1o60MMv7zij)

2\. Click **Save**.

### 3. Connect Activity Data

Activity data is used to calculate popularity and top users. To enable event logs in select star you will need to [setup AWS CloudTrail](/integrations/quicksight/event-logs), once complete you should have the following information available:

* **Event Logs Bucket** - The name of the bucket created in the S3.
* **Event Logs Bucket Prefix** - The prefix specified for your CloudTrail logs.

![Connect Activity Data](/files/1g9JOvFT1FsuK8euToFs)

3\. When the sync is complete, you'll be able to explore QuickSight in Select Star.


# Event Logs

Available event logs sources for QuickSight data sources:

* AWS CloudTrail

## Setup CloudTrail logs on a S3 bucket

Setting up AWS CloudTrail to store data in an S3 bucket for tracking API usage in AWS QuickSight involves several steps. Here’s a step-by-step guide using the AWS Management Console:

### Step 1: Sign in to the AWS Management Console

1. Go to the [AWS Management Console](https://aws.amazon.com/console/).
2. Sign in using your AWS account credentials.

### Step 2: Create an S3 Bucket to Store CloudTrail Logs

1. **Navigate to the S3 Service**:
   * In the AWS Management Console, search for "S3" in the search bar at the top, and select **S3** from the results.
2. **Create a New Bucket**:
   * Click on **Create bucket**.
   * Provide a **unique bucket name** (e.g., `cloudtrail-logs-youraccountname`).
   * Choose the **AWS Region** where you want to create the bucket.
   * Click **Create bucket**.

### Step 3: Create a CloudTrail to Monitor API Calls

1. **Navigate to the CloudTrail Service**:
   * In the AWS Management Console, search for "CloudTrail" in the search bar, and select **CloudTrail** from the results.
2. **Create a Trail**:
   * In the CloudTrail dashboard, click on **Create trail**.
   * Step 1:
     * **Trail name**: Enter a name for your trail (e.g., `QuickSightAPILogging`).
     * In the **Storage location** section, under **Create a new S3 bucket or use an existing one**, select **Existing S3 bucket**.
     * **S3 bucket**: Choose the S3 bucket you created earlier (e.g., `cloudtrail-logs-youraccountname`).
     * Prefix: Optionally, specify a prefix for your CloudTrail logs (e.g., `quicksight-logs/`). Take note of this value, it will be required later for your data source configuration (e.g., `quicksight-logs/AWSLogs/792169733636`)
   * Step 2:
     * **Management events**: Enable management events if not already enabled. This logs control plane activities (e.g., CreateTable in QuickSight).
     * **Management events**: Choose the Read option.
   * Step 3:
     * Review and Create

### Step 4: Verify CloudTrail is Logging

1. **Check CloudTrail Logs**:
   * Return to the CloudTrail dashboard.
   * View the **Events history** to confirm that CloudTrail is logging events. Open Dashboards and Analyses to generate GetDashboard and GetAnalysis events. (It may take seconds or even a minute between opening a Dashboard and event generation on CloudTrail).
2. **Check S3 Bucket**:
   * Navigate to the S3 service and open your bucket.
   * Verify that logs are being delivered to the specified folder in the bucket.

### Step 5: Query CloudTrail Logs for AWS QuickSight API Usage

1. **Identify QuickSight API Calls**:
   * Look for API calls related to QuickSight, such as GetDashboard, GetAnalysis (This may take a while to show up too).
2. **Download and Analyze Logs**:
   * You can download the logs from your S3 bucket and analyze them manually or using tools like Amazon Athena to query logs directly in S3.

### Step 6: The information needed for Select Star

* Event Log Bucket: The bucket name you chose on step 2.2 (e.g., `cloudtrail-logs-youraccountname`).
* Event Log Bucket Prefix: The prefix name chosen on step 3.2, plus `CloudTrail/yourregion` (e.g., `/AWSLogs/792169733636/CloudTrail/us-east-2/`)

### Conclusion

By following these steps, you’ll have AWS CloudTrail configured to log API calls, store them in an S3 bucket, and track AWS QuickSight usage. You can use these logs for security audits, compliance, or understanding how QuickSight is used in your organization.

## Setup the access so Select Star can read the S3 stored logs

Updating an existing AWS IAM role to enable access to an S3 bucket containing CloudTrail logs involves attaching an appropriate policy to the IAM role. Here’s a step-by-step guide:

### Step 1: Sign in to the AWS Management Console

1. **Go to the AWS Management Console**: Navigate to <https://aws.amazon.com/>.
2. **Sign in**: Enter your credentials to log in.

### Step 2: Navigate to the IAM Service

1. **Search for IAM**:
   * In the AWS Management Console, use the search bar at the top to search for "IAM" and select **IAM** from the results.
2. **Open IAM Roles**:
   * In the IAM dashboard, click on **Roles** in the left-hand navigation pane.

### Step 3: Locate the IAM Role (`CrossAccountQuicksight`)

1. **Search for the Role**:
   * Use the search bar to find the IAM role named `CrossAccountQuicksight`.
2. **Select the Role**:
   * Click on the role name `CrossAccountQuicksight` to open its configuration page.

### Step 4: Attach a Policy to the IAM Role to Access the S3 Bucket

1. **Click on Add Permissions**:
   * On the role's page, click the **Add permissions** button.
   * Choose **Create inline policy** from the dropdown menu.
2. **Create a Custom Policy**:
   * Since you need to grant access to a specific S3 bucket, you’ll create a custom inline policy.
   * Click on the **JSON** tab to enter the policy directly.
3. **Enter the S3 Bucket Policy**:

   * Add actions `s3:GetObject` and `s3:ListBucket`.
   * Add the `Event Log Bucket` arn and the `Event Log Bucket Prefix` folder arn (/\*) as Resources.
   * Here's a sample policy:

   ```json
   {
       "Version": "2012-10-17",
       "Statement": [
           {
               "Effect": "Allow",
               "Action": [
                   "s3:GetObject",
                   "s3:ListBucket"
               ],
               "Resource": [
                   "arn:aws:s3:::cloudtrail-logs-youraccountname",
                   "arn:aws:s3:::cloudtrail-logs-youraccountname/AWSLogs/792169733636/CloudTrail/us-east-2/*"
               ]
           }
       ]
   }
   ```
4. **Review and Attach the Policy**:
   * After entering the policy, click **Next**.
   * Provide a name for the policy (e.g., `SelectStarAccessS3CloudTrailLogs`).
   * Review the details and click **Create policy** to attach it to the role.

### Step 5: Verify the Role Update

1. **Review Attached Policies**:
   * Back on the IAM role page, review the list of attached policies to ensure that your new policy (`AccessS3CloudTrailLogs`) is listed.
2. **Test the Role** (Optional):
   * If you have access to the system where this role is used, you can test it by attempting to access the CloudTrail logs in the specified S3 bucket.

### Step 6: Save and Exit

1. **Save the Configuration**:
   * Ensure that all changes are saved and the policy is properly attached.
2. **Exit the IAM Console**:
   * You can now exit the IAM console.

### Conclusion

By following these steps, you’ve successfully updated the IAM role `CrossAccountQuicksight` to enable it to access the S3 bucket where your CloudTrail logs are stored. This ensures that any service or application using this role can retrieve and process the CloudTrail logs as needed.


# Hex (beta)

To connect Hex to Select Star, you will need to...

1. [Create an API key in Hex](#id-1.-create-an-api-key-in-hex)
2. [Create a new datasource in Select Star](#id-2.-create-a-new-data-source-in-select-star)

{% hint style="info" %}
The integration with Hex currently only supports project level lineage and does not support chart level lineage. This is due to a limitation in the metadata provided via Hex APIs.
{% endhint %}

## 1. Create an API Key in Hex

1. Go to Hex and sign in with your account
2. Once logged in, click on settings icon (at the bottom left of the screen)
3. Under **Account,** select **API keys**
4. Create new **Workspace token**:

   * provide a meaningful description
   * set an appropriate duration for the expiration date (recommend `no expiration`)
   * set the API Scope to `Read projects`

   ![](/files/jZZ4GG2J4N0nq35Kkmfy)
5. Once the key is generated, copy it to use to connect to Hex in Select Star.

## 2. Create a New Data Source in Select Star

* **Display Name** - The name you want to give to your new data source.
* **API Key** - The workspace token created in Hex from step 1.
* **Hex URL** - The URL copied from your hex instance, such as `https://app.hex.tech/selectstar/`, composed of two key parts:
  * **Base URL** - The primary domain of your Hex instance (e.g., `https://app.hex.tech`).
  * **Site ID** - The unique identifier for your specific workspace or site (e.g., `selectstar`)

![Create Data Source](/files/fbITXSEBh0JMDz3tjp7u)

Your metadata should start loading automatically, and once the sync is complete you'll be able to see Hex projects and their lineage in Select Star.


# Omni (beta)

Follow these steps to connect your Omni instance to Select Star.

* Generate an API Key
* Connect Omni to Select Star

1. ## Generate an API Key

You'll need an Admin account to generate API Keys. Once logged in, go to:

* Settings
* API Access
* Organization keys (used for service integrations)&#x20;
* Generate a New Key

2. ## Connect Omni to Select Star

* **Display name:** The name you want to give to your new data source
* **Domain:** Your domain on Omni, it's available on your access URL. E.g.: https\://**selectstar**.omniapp.co (**selectstar** in this case)
* **API Key:** The organization API Key you have just generated

<figure><img src="/files/mAlzW6TlARwzaSNxk9hF" alt=""><figcaption></figcaption></figure>


# Slack

Select Star now integrates with Slack, making it easier than ever to find and share links to data assets. Learn how this integration can improve your team's productivity and help you stay organized.

Select Star provides a powerful application for Slack for a number of uses:

* Users can quickly find and share links to data assets with other people in the organization
* The `@Select Star` bot can be added to public and private channels to answer data questions based on the information in Select Star.

Using the Select Star App for Slack, you can leverage the power of Select Star's search to quickly send a starting point for others to explore, receive Select Star notifications, or browse through search results.

## Video Tutorial

{% embed url="<https://www.loom.com/share/f381c5b5bf2d4e52b519d9c9bf7897fb?hideEmbedTopBar=true>" %}

## Installation

### Prerequisites

You will need admin access to your organization's Slack.

### Instructions

Visit <https://slack.selectstar.com/> and click **Add to Slack.** You'll be redirected Slack in the browser and asked to grant permissions.

Click **Allow**. You'll be guided through a few simple integration steps by Slack. Then you'll be all set up to start using the app!

<figure><img src="/files/4ULAt9cyNljp2min5lWb" alt=""><figcaption></figcaption></figure>

### Re-Installation

If you need to re-install the Select Star App for Slack (such as if you are upgrading to leverage the new AI based capabilities), you can re-install from the following link: <https://slack.production.selectstar.com/slack/install>.

## Connect a User Account

In order to leverage the Select Star App features such as AI powered chat, notifications, and search, your user account must be connected.\
\
Head to Slack and find Select Star in the list of Apps.

<figure><img src="/files/iXdiVXeqrMzARDAXV4M4" alt=""><figcaption></figcaption></figure>

Under the Home tab you will see a button to **Connect an account**. This will link your Slack account with your Select Star user account so that you can receive notifications and search in Slack.

<figure><img src="/files/KSjAyYEBdq4JR7VMZKNo" alt=""><figcaption></figcaption></figure>

Alternatively, clicking **Link User Account** under **User Settings** will take you through the same steps.

<figure><img src="/files/YSoc0HuwPXJ1JneTV60W" alt=""><figcaption><p>User settings to link the user's Slack account to their Select Star account</p></figcaption></figure>

At least one user must connect their account to use the AI powered chat features.

## Adding Select Star to a Slack Channel

To use Ask AI Chatbot and schema change notifications in Slack, you first need to add Select Star to the channel where you want to use these features.

1\. Open Slack and navigate to the channel where you want to receive notifications.

2\. Type /invite **@Select Star** in the message input and press Enter.

3\. Select Star will be added to the channel, and notifications will be enabled.

## Ask AI - Answer Data Questions

{% hint style="info" %}
AI powered chat in Slack needs to be separately enabled for your account. If you'd like to enable this feature, please reach out to <support@getselectstar.com>.
{% endhint %}

Select Star's Slack bot can help detect and identify data questions that can be answered by information you have in Select Star. Relevant questions include:

* Learning more about your data
* Finding the right content
* Generating queries

[Read more on the AI chatbot](https://docs.selectstar.com/features/ask-ai-chatbot#ai-chatbot).

#### How to use the AI chatbot in Slack

* The Select Star bot `@Select Star` must be added to the channel (public or private)
* The bot will identify questions that it can help answer, and will automatically ask to answer
* Users can also mention `@Select Star` directly in a message or thread to get a response

{% hint style="info" %}
**Important note**: once setup, any user can use the Select Star AI chatbot to answer questions. They do not need a Select Star account, and will be able to access all data that is connected and searchable within Select Star.
{% endhint %}

## Slack Notifications

### User Notifications

When a users add the Select Star App to their Slack instance, they will start receiving notifications when changes are made in Select Star. These are the same notifications they would see in the Select Star application.

You'll receive a notification for:

* Metadata changes
  * You are added as an [Owner](/data-management/data-ownership) to a data asset.
  * Someone [comments](/features/discussion) on a data asset you own.
  * Someone modifies the description or tags on a data asset you own
* Schema changes
  * There is a schema change in an object that you own.
* Mentions
  * Someone replies to a comment you posted, or mentions you in a comment.

Users can configure their notification settings on the [User Settings](/user-management/account-and-user-settings#user-settings) page.

### Team Notifications

Similar to user notifications, admins can configure team notifications to send to a Slack channel. See [Team Notifications](/features/teams#team-notifications).

1\. Go to Admin > Teams > Edit in Select Star.

2\. You’ll see a list of Slack channels where Select Star has been added.

3\. Select the channels or email addresswhere you want to receive notifications.

4\. Save your settings.

<figure><img src="/files/QE48vNg35WBx9azPAbrA" alt=""><figcaption></figcaption></figure>

### Organization Notifications

Admins can configure organization wide notifications for schema changes. Admins can select a Slack channel to receive notifications for schema changes for:

* Assets with business or technical owners, OR
* All assets

<figure><img src="/files/6Q2I4QlrSOZhBNA6Ew95" alt=""><figcaption><p>Configure organization Slack schema changes</p></figcaption></figure>

## Slack Commands

Currently, the app has two commands.

* [**Help Command**](#help-command): Learn the usage of the Select Star App for Slack.
* [**Search Command**](#search-command): Search and post results from Select Star in Slack. You can also specify the type of search results you'd like to see.

### Help Command

You can type `/selectstar-help` into Slack at any time to see information on how to use the app.

<figure><img src="/files/Em8PJISmXQhM711hDBtk" alt=""><figcaption></figcaption></figure>

### Search Command

[Search](/features/search) Select Star using the command `/selectstar <search query>` or `/s* <search query>`. Both commands work the same way; one is simply shorter than the other.

For example, if you search `/selectstar customers` in Select Star, you may see something like this:

<figure><img src="/files/TJujN0LmALgDi4zi4kWn" alt=""><figcaption></figcaption></figure>

You'll see the results in the same order you would in Select Star, where the most relevant and popular items show first.

If the first result isn't what you're looking for, you can click the `Next` and `Previous` buttons to browse through the results, or choose to `See all results in Select Star` to open the [search results page](/features/search#search-results-page) in your browser.

Search results from the Select Star App for Slack will be visible only to you, unless you choose to click the `Post` button, which will share the results to everyone in the channel.

<figure><img src="/files/erz1wEjifPFBAkNra6jg" alt=""><figcaption></figcaption></figure>

#### Searching by Type

If you are looking for a specific kind of data, such as a database table or Tableau view, you can specify the type in the command.

Based on the same example above, if you wanted to find database tables with "Customer" in the name, you would search `/selectstar customer type:columns`, with no space between `type:` and `columns`.

<figure><img src="/files/9EpbZSV5J0dFo4ghrwRi" alt=""><figcaption></figcaption></figure>

The list of data types you can filter by:

* **Database**: table, column, user, schema, database
* **Looker**: looker\_explore, looker\_explore\_field, looker\_dashboard
* **Mode**: mode\_space, mode\_report
* **Tableau**: tableau\_view, tableau\_field, tableau\_data\_source
* **Power** BI: power\_bi\_folder, power\_bi\_dashboard. power\_bi\_report
* **Sigma**: sigma\_dashboards, sigma\_folders
* **Thoughtspot**: thoughtspot\_liveboards, thoughtspot\_answers, thoughtspot\_tables, thoughtspot\_views, thoughtspot\_worksheets, thoughtspot\_columns
* **QuickSight**: quicksight\_dashboards, quicksight\_reports, quicksight\_bidashboardelements, quicksight\_bidatasets, quicksight\_bicolumns

### Slack Settings

Under **Admin > Apps > Slack** you can see which channels the Select Star Slackbot has been added to and configure unfurl settings for Select Star URLs.

<figure><img src="/files/SJ3kcqUuhrTRtEdlM89X" alt=""><figcaption><p>Slack integration settings page</p></figcaption></figure>

## Removing Slackbot from a Channel

To remove the Select Star slackbot from a specific channel use the `/remove` command in the channel you want the bot to leave.


# Monte Carlo

Get visibility into your data quality issues directly from Select Star.

{% hint style="info" %}
The Monte Carlo integration is in beta. Please reach out to <mark style="color:blue;"><support@getselectstar.com></mark> to get started.
{% endhint %}

The Monte Carlo integration helps surface data quality information into Select Star, the place most users go to find and discovery datasets and get more context. This integration helps spread awareness of data quality and guides decisions on what data to use.

The integration surfaces data quality information from Monte Carlo monitors directly on the tables themselves.

Note: only monitors that are enabled and active are shown. Any monitors that are disabled, snoozed, in training, have insufficient data, misconfigured or otherwise unable to run are not displayed.

<figure><img src="/files/gUtBF2msuTys5VIsRSFA" alt=""><figcaption><p>Monte Carlo monitor status on the Select Star table page</p></figcaption></figure>

## Getting Connected

### 1. Generate an Account Service Key in Monte Carlo

To create an account-service API key for connecting to Monte Carlo:

1. In Monte Carlo, go to **Settings** in the top header.
2. In the left menu click **API** and then click **Account Service Keys** at the top.
3. From the *Account Service Keys* page, click **Create Key** and enter the following:
   1. In *Description*, add a useful description for your API key — i.e. `Select Star`.
   2. For the *Authorization Groups* dropdown, select **Editors (All)** to provide [minimum permissions](https://docs.getmontecarlo.com/docs/authorization#managed-roles-and-groups) for crawling Monte Carlo.
   3. For *Expires After*, keep the default selection or select a preferred option.
   4. Click **Create** to finish creating the account-service API key.
4. Copy the *Key ID* and *Secret* and store them securely.

### 2. Configure the Integration in Select Star

In Select Star, configure Monte Carlo as an **App** to set up the integration. Note - this must be done by an admin.

1. Click **your profile** in the top right, and go to **Settings**.
2. Go to **Admin > Apps** and on the Monte Carlo tile, click **Connect.**
3. Enter the **API Key** and **Secret** and click **Save**.

It will take a few minutes to ingest the monitor information from Monte Carlo.\\

### Integration Details

Once a day Select Star will ingest all monitors and latest statuses for any alerts. Select Star uses webhooks that will appear as **Audiences** within Monte Carlo to be able to get fast updates when an alert is created or status has changed.

Select Star maps the Monte Carlo Alert Status to a status for the Monitor, so you can quickly see where your data issues are.

| Monte Carlo Alert Status | Select Star Monitor Status |
| ------------------------ | -------------------------- |
| No Status                | ❌ Failed                   |
| Investigating            | ❌ Failed                   |
| Work in Progress         | ❌ Failed                   |
| Fixed                    | ✅ Passed                   |
| Expected                 | ✅ Passed                   |
| No Action Needed         | ✅ Passed                   |
| False Positive           | ✅ Passed                   |


# Private Network

For customers hosting data sources on private networks, such AWS VPC, or on-premises, our platform supports several secure and flexible integration methods. These methods are designed to ensure seamless data integration while upholding the highest security standards.

**Before beginning any integration process of data sources in private network, please contact our technical support team.** Our team will work with you to review your specific requirements and environment, helping to select the most effective solution.

Here’s how you can connect your data sources in private network to our platform:

## 1. Using Load Balancers

Customers can configure a load balancer to expose their data sources securely to us. In this case, instead of a direct connection to the data source, a connection via loadbalancer is used. The loadbalancer address should be indicated as the data source's hostname configuration.

Steps to implement:

* **Setup loadbalancer**: Customer to setup loadbalancer routing to the data source.
* **Connection Setup**: Our technical support to ensure that load balancer is used to connect to your data source

**For AWS Users:** Utilize an AWS Network Load Balancer (NLB) along with [security groups](https://aws.amazon.com/about-aws/whats-new/2023/08/network-load-balancer-supports-security-groups/) to manage access controls effectively.

**Security Measures**: The loadbalancer must be publicly accessible from the Internet. We recommend implementing IP filtering to restrict access exclusively to our platform, enhancing your security posture.

## 2. Using AWS PrivateLink

For customers using AWS, AWS PrivateLink provides a secure and scalable way to connect services across different accounts and VPCs without exposing data to the public internet. This solution does not require further maintenance or updates after one-time setup.

Steps to implement:

* **Initial contact**: Contact us so our support team can provide the necessary AWS Account ID for your environment.
* **Setup AWS PrivateLink Service**: Create an AWS PrivateLink Service in your VPC, authorize the AWS region `us-east-2` region and our AWS Account. Share with our team the service name and your AWS region.
* **Setup AWS PrivateLink Endpoint**: Our team will set up an AWS PrivateLink Endpoint in our VPC that connects to your AWS PrivateLink Service.
* **Connection Setup**: Our technical support will ensure the AWS PrivateLink is correctly established and used to connect to your data source securely.

Depending on the data source and environment, it may also require other cloud resources, such as internal AWS NLB, AWS Lambda.

Helpful links:

* [Working with Redshift-managed VPC endpoints](https://docs.aws.amazon.com/redshift/latest/mgmt/managing-cluster-cross-vpc.html)
* [Access Amazon RDS across VPCs using AWS PrivateLink and Network Load Balancer](https://aws.amazon.com/blogs/database/access-amazon-rds-across-vpcs-using-aws-privatelink-and-network-load-balancer/)
* [AWS PrivateLink now Supports Access Over VPC Peering](https://aws.amazon.com/about-aws/whats-new/2019/03/aws-privatelink-now-supports-access-over-vpc-peering/)
* [Introducing Cross-Region Connectivity for AWS PrivateLink](https://aws.amazon.com/blogs/networking-and-content-delivery/introducing-cross-region-connectivity-for-aws-privatelink/)

## 3. Using SSH Tunneling

SSH tunneling can be setup to securely connect to data sources within a customer's private network. This solution allows for an encrypted point-to-point connection via the Internet.

Steps to implement:

* **Public Key Provision**: Our team will provide a public key specifically for the data source, allowing secure connection to the SSH bastion.
* **SSH Bastion Setup**: Setup an SSH bastion host that can access the data source directly.
* **Access Authorization**: Grant our platform access to an SSH bastion host within your network.
* **Connection Setup**: Our team will ensure that the SSH tunnel is correctly established and used to securely connect to your data source.

**Security Measures**: The SSH bastion must be publicly accessible from the Internet. We recommend implementing IP filtering to restrict access exclusively to our platform, enhancing your security posture.

## 4. Using Reverse Tunneling

If establishing a direct SSH connection is not feasible, reverse tunneling offers a secure alternative by using an external SSH broker to handle connections.

This method secures your data transmission by establishing an intermediary that handles all external connections, thereby not exposing any host in your private network directly to the Internet.

Steps to implement:

* **Public Key Provision**: Our team will provide a public key specifically for the data source, allowing secure connection to the SSH broker.
* **SSH Broker Setup**: Establish dedicated hosts that will act as SSH brokers. This could be an EC2 instance configured solely with an SSH server, or any other suitable VPS.
* **Tunnel Creation**: Set up an SSH or VPN tunnel from your private network to the SSH broker, so SSH broker can access data source.
* **Access Authorization**: Grant our platform access to an SSH broker host.
* **Connection Setup**: Our team will ensure that the SSH tunnel is correctly established and used to securely connect to your data source.

**Security Measures**: The SSH broker must be publicly accessible from the Internet. We recommend implementing IP filtering to restrict access exclusively to our platform, enhancing your security posture.

## 5. Other methods

We are open to meeting the needs of rigorous environments. If none of the solutions presented meet your requirements, still contact us to explore less common solutions.

## Support and Troubleshooting

If you encounter any issues or require assistance during the setup process, our dedicated support team is available to help you with troubleshooting and guidance to ensure a smooth integration process.


# Features

Select Star is a data exploration platform that provides powerful tools for analyzing and visualizing data. Learn more about its key features and how they can help you make the most of your data.

This section contains the following articles:

{% content-ref url="/pages/-MgC7Lp0AFs\_b7oMHoZ0" %}
[Search](/features/search)
{% endcontent-ref %}

{% content-ref url="/pages/-MiXUMWODT7Itc685YFN" %}
[Table Page](/features/table-page)
{% endcontent-ref %}

{% content-ref url="/pages/cpWHwlprfqVhIMUFMWKk" %}
[Database Page](/features/database-page)
{% endcontent-ref %}

{% content-ref url="/pages/-MiTnqOjij5qhcKrQ3XW" %}
[Dashboard Page](/features/dashboard-page)
{% endcontent-ref %}

{% content-ref url="/pages/-MgBz1w\_cd0EvhqPLll0" %}
[Data Lineage](/features/lineage)
{% endcontent-ref %}

{% content-ref url="/pages/-MgC7SVl\_fmOS7RRRZWX" %}
[Queries & Joins](/features/queries-and-joins)
{% endcontent-ref %}

{% content-ref url="/pages/-MgC83u5A0eoWXkwcKck" %}
[Tags](/features/tags)
{% endcontent-ref %}

{% content-ref url="/pages/-MgC7wTcAW82aXgKjVzK" %}
[Discussion](/features/discussion)
{% endcontent-ref %}

{% content-ref url="/pages/PhqASr9qNXa7XR3Iclno" %}
[Automated Documentation](/features/auto-documentation)
{% endcontent-ref %}

{% content-ref url="/pages/-MgIK5AM1-orNYsHjcEF" %}
[Metrics](/features/documents/metrics)
{% endcontent-ref %}

{% content-ref url="/pages/qwJWKKirRiY5zKhjCtke" %}
[MCP Server](/features/mcp-server)
{% endcontent-ref %}

{% content-ref url="/pages/KR1ECZxYKsO741OUr0pq" %}
[Semantic Models](/features/semantic-models)
{% endcontent-ref %}


# Search

Utilize Select Star's powerful search feature to find various data objects, even if you don't have their exact names. Explore tables, columns, reports and more.

Select Star’s Search is designed to help you find any data objects (tables, columns, Tableau Workbooks, Mode Reports, Looker Explores, etc.) without even knowing exactly what they’re called.

Use search in Select Star to answer questions such as...

* [Where's my data?](/data-discovery/wheres-my-data)
* [Where's my dashboard?](/data-discovery/wheres-my-dashboard)

Watch this video to learn more about Select Star's universal search, or keep reading for more details:

### Search Types

| Type                                                                                                  | Description                                                                                                                                                                      |
| ----------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Keyword Search**                                                                                    | Traditional token-based search that matches specific keywords or phrases entered by users, used in the **Top Search**.                                                           |
| **Natural Language Search (Requires enabling** [AI Settings](/features/ai-features-and-settings)**)** | Allows users to input queries in a conversational manner, as they would ask a question to another person in conversational English and show them on the **Search Results Page**. |

## Top Search

The search bar at the top of the page is always available to find or quickly navigate to any data asset.

<figure><img src="/files/hHhbaJ144EweCDICyHmK" alt=""><figcaption></figcaption></figure>

Space works as a delimiter in the search, so if you search for `order items` you'll get results with both "order" *and* "items", whereas if you search for `order_items` you'll get results that are looking for "order\_items" as one word.

Note that when you search for multiple words, the order of the search terms matters!

<figure><img src="/files/XemL3pc5Yk3dlPFk7Vho" alt=""><figcaption><p>Search results for "order" and "items", separated by a space.</p></figcaption></figure>

<figure><img src="/files/pLOKP1WrnYzpgTh7nyfk" alt=""><figcaption><p>Search results for "order_items", treated as one word.</p></figcaption></figure>

All search results are sorted by relevance and popularity, so you’ll see what’s being used the most on the top. You won't have to worry about someone's [test tables](/data-source-management/manage-data-sources#hide-datasets) or dashboards showing up at the top of the search.

Click on different tabs below the search bar to see results from a specific data source. For instance if you're looking for a Looker Explore, click the **Looker** tab and then select **Looker Explore**.

<figure><img src="/files/AJy0sAsGE7mvRUjjeuYN" alt=""><figcaption></figcaption></figure>

## Search Results Page

After you've typed a query into the top search, hit `Enter`/`Return` on your keyboard to see a full search results page and browse more high level information at once.

The tabs available in the top search dropdown are separated into individual filters on the side of the page, under **Asset Types**.

<figure><img src="/files/hDnBFHXU4SkXLSFPsn0Z" alt=""><figcaption></figcaption></figure>

From this view, you can also choose to display or hide other items, such as [Tags](/features/tags) and [Owners](/data-management/data-ownership), by adjusting the column filters.

<figure><img src="/files/b9tCT3GRzMm5tXDJKMmC" alt=""><figcaption></figcaption></figure>

### Natural Language Search

NLS lets users search for data assets by asking questions in everyday language - no need to know specific keywords or table names. Moreover, we offer multi-language support. On the other hand, if a query isn’t classified as a natural language question, it automatically falls back to standard keyword-based search using Elasticsearch.

Alongside the search results, NLS displays an AI-generated response similar to a chatbot, summarizing relevant information from the matched data assets. Users can click a “Continue the Conversation” button to ask follow-up questions, enabling a more interactive and conversational search experience.

<figure><img src="/files/WWDg7iGwpBPjDA3T3lUH" alt=""><figcaption></figcaption></figure>

### **Example Queries**

* How is revenue calculated?
* Where can I find shipping facility preferences?
* Can you find any dashboards or tables that mention revenue?
* What information do we have on retention?
* Who can answer questions about the business segments dashboard?

### Enabling Natural Language Search

1. Make sure the **"Enable AI Features"** setting is turned on in the **AI Settings** section of the Admin panel ([AI Settings](/features/ai-features-and-settings)).
2. Once enabled, NLS will be active on the main **Search page**. The top navigation search bar will continue to use keyword-based search.
3. Note: We still can disable this specific feature upon request for your organization. Contact your **Select Star account manager** or support team to request it.

## Search Ranking

<table><thead><tr><th width="390.41015625">Keyword Search</th><th>Natural Language Search</th></tr></thead><tbody><tr><td><p>Select Star ranks search results on a number of factors, including:</p><ul><li>match type (name, breadcrumb, description)</li><li>popularity</li><li>asset type</li><li>tag</li></ul></td><td>Search results are ranked based on several factors such as match type, asset popularity, and tags. For NLS queries, results are primarily ranked by the AI model’s relevance scoring, but users can still apply sorting and filters as usual.</td></tr></tbody></table>

### Tag Boost in Keyword Search

Customers put a lot of effort to tag and organize data, and we want that to work for you. That is why we have updated our search algorithm to provide a slight boost for data assets tagged with specific tags. Data Assets with a tags with the following icon will receive a positive boost in the search results:

<figure><img src="/files/F0GCPnueoSyyxQVMisPh" alt=""><figcaption><p>Tag icons with positive search boost</p></figcaption></figure>


# Table Page

Learn about basic and operational information, and navigate through different table tabs for comprehensive insights.

Database tables in Select Star each have their own page where you can explore a variety of information.

{% embed url="<https://www.loom.com/share/5f3cfb28b2db48fba0a6f23276b060a5?hideEmbedTopBar=true>" %}

This page gives an overview of the information available in three different areas:

* [Basic Information](#basic-information)
* [Operational Information](#operational-information)
* [Table Tabs](#table-tabs)

Click the links within this page to learn more about specific features.

<div data-full-width="true"><figure><img src="/files/YgaLVaTqJpok0X2uX3FN" alt=""><figcaption></figcaption></figure></div>

## Basic Information

<div data-full-width="true"><figure><img src="/files/4WFLPK82qLoJekhLEUB4" alt=""><figcaption></figcaption></figure></div>

* The full table path of the table in the database.
* [**Tags**](/features/tags) that have been applied to the table.
* A [**description**](/data-management/add-documentation) of the table. Descriptions can synced from Snowflake or dbt's [persist\_docs](https://docs.getdbt.com/reference/resource-configs/persist_docs) function, or you can add them directly in Select Star.

## Operational Information

<div data-full-width="true"><figure><img src="/files/mYHkx7PRFSRn0ccxpRCO" alt=""><figcaption></figcaption></figure></div>

* [**Business and Technical Owners**](/data-management/data-ownership) - Subject matter experts for this particular table, capable of answering business and technical questions about the data. These can be assigned by [Admins and Data Managers](/user-management/user-roles).
* **Popularity** - Calculated by Select Star. This is the # of times queried by # of users over a customizable period of 30-90 days.
* **Table Size** - The size of the table and number of rows.
* **Last Update** - When the table metadata was last updated. This value does **not** indicate when the *data values* were last updated.
* **Last Refreshed** - Indicates the most recent time a table was updated based on metadata and query activity. This helps you understand how up-to-date a table is before using it for analysis or reporting. If a table hasn’t been refreshed in a while, it may indicate stale data.

  How it works:

  * For most tables, this reflects the last time the table was modified in your data warehouse, such as when new data was added or structure changes were made.Comment
  * If there’s no direct record of when the table was last modified, Select Star will use the most recent metadata sync timestamp as the Last Refreshed date.
  * The Last Refreshed date is always based on the latest available information and ensures consistency across your datasets.

  Note: Some refresh delays may occur due to ingestion schedules or metadata sync timing.
* **\<SQL>** - When available, clicking this will show the SQL used to create the table or view.

#### Video Tutorial: Popularity

{% embed url="<https://www.loom.com/share/c63f6380dac84e3abb4b9a0f027c1ff5?hideEmbedTopBar=true>" %}

## Table Tabs

<div data-full-width="true"><figure><img src="/files/NihjF2I7DLi4knGNfejX" alt=""><figcaption><p>Table tabs in Select Star</p></figcaption></figure></div>

Tabs will give you special information about the table data and [how the data is used](/data-discovery/how-can-i-use-this-data) in the context of your organization.

* **Columns** - The columns in that table, along with any [descriptions](/data-management/add-documentation) or [Tags](/features/tags) they have. See how popular individual columns in a table are, as well as which columns are used to join to other tables. Columns can be ordered three ways on a table page. A -> Z, Z -> A, and Ordinal. Click the header to change the order.
* [**Lineage**](/features/lineage) - This tab will have an interactive graph to explore of Upstream Sources and Downstream Targets connected to this table, as well all the Downstream Dashboards.
* [**Queries & Joins**](/features/queries-and-joins) - Recent and popular queries and join conditions you can use to understand [how the data is used](/data-discovery/how-can-i-use-this-data).
* **Related** - This tab has a list of Similar and Related Tables, where Similar Tables are tables that have the exact same table structure as this table, while Related Tables are tables that can be joined to this table.
* **Top Users** - The people or [service accounts](/getting-started/mark-service-accounts) who are most relevant for this table. This is separated into Table Users who have the most queries against this table, and Dashboard Users who have the most views/queries of directly downstream dashboards.
* **Data Quality -** If you have [Monte Carlo](/integrations/monte-carlo) connected as a datasource, or a dbt data source set up, including dbt tests, you will see any monitors or tests that are related to the table. If all tests have passed you will see a green checkmark on the data quality tab, otherwise the red dot will indicate how many tests / monitors have failed.
* **Preview** - A sample of data values in this table. This is disabled by default, as Select Star only looks at metadata from your sources. Contact a Select Star representative if you would like to enable this by sending some sample data values.
  * Note: Preview data is not persisted, it is queried each time it is requested.
* [**Discussion**](/features/discussion) - Anyone with a Select Star account can post [questions](/data-discovery/i-have-a-data-question) or comments on this dataset, which are then searchable.


# Database Page

In Select Star, the database page provides an overview of your database, including tables and schemas. Explore and manage your data insights in a single platform.

You can see an overview of your database with all tables and schemas in Select Star.

Learn more about the database page by watching this video:

{% embed url="<https://www.loom.com/share/5bb95501cb7a4a3a9de38f71ebb3efab?hideEmbedTopBar=true>" %}

Read about database pages by visiting the links below.

{% content-ref url="/pages/-MgSY-bV9bXpwypJdJVv" %}
[Getting Started: Snowflake](/learning-data/getting-started-snowflake)
{% endcontent-ref %}

{% content-ref url="/pages/-MgSY3JsJz5uuNb23QXn" %}
[Getting Started: BigQuery](/learning-data/getting-started-bigquery)
{% endcontent-ref %}


# Dashboard Page

Explore dashboards, reports, and looks in Select Star with dedicated pages. Access basic and operational information, dashboard tabs, and engage in discussions.

Dashboards, reports and looks in Select Star each have their own page where you can explore a variety of information.

{% hint style="info" %}
Different [BI Tools](/integrations#bi-tool-integrations) use different words to describe dashboards, reports, or other specific views they use to organize data. These docs try to use language specific to the tool when possible, but in general this type of content is referred to as "dashboards."
{% endhint %}

This page gives a brief overview of the information available in three different areas:

* [Basic Information](#basic-information)
* [Operational Information](#operational-information)
* [Dashboard Tabs](#dashboard-tabs)

Click the links within this page to learn more about a specific feature.

<figure><img src="/files/mTspwA0ZbyDhNBFPy8yB" alt=""><figcaption></figcaption></figure>

## Basic Information

<figure><img src="/files/GaAb7ARDE1lGp9FAlftR" alt=""><figcaption></figcaption></figure>

* The name of the dashboard or report.
* A button to **Open** the dashboard in your BI tool. This will only work if you have an account for the BI tool and are logged in.
* [**Tags**](/features/tags) that have been applied to the dashboard.

## Operational Information

* A button to **Open** the dashboard in your BI tool. This will only work if you have an account for the BI tool and are logged in.
* **Created By** - \*\*\*\* The BI Tool user who created the dashboard. If this user has created an account in Select Star with the same email, their different accounts will be listed under their [User Profile](/data-discovery/im-new-to-the-team#team-profiles).
* **Popularity** - Calculated by Select Star. This is based on impressions. For dashboards it is the # of views over a period of 30-90 days.
* **Loading Status** - This shows whether or not the dashboard has run successfully.
  * :white\_check\_mark: Indicates the dashboard is loading successfully
  * ​:question: Indicates uncertain loading status. This dashboard has not been run in a while
  * :warning: This dashboard's last run was unsuccessful. One or more of the queries has an error.
* **Last Edited** - When the metadata of the dashboard was last changed in the BI Tool.
* **Last Run** - When the dashboard was last run. This is currently updated during Select Star's daily metadata sync. If you run your dashboards more than once a day, this may not show the most recent run.

## Dashboard Tabs

<figure><img src="/files/yrWAFFHHQ7eRmCv4LVmA" alt=""><figcaption></figcaption></figure>

Tabs will give you special information about how the dashboard is used and how it connects to other data in your organization.

* **Overview Tab -** The overview tab gives all the operational information for the dashboard such as
  * The [**description**](/data-management/add-documentation) of the dashboard.
  * **Created By** - The BI Tool user who created the dashboard. If this user has created an account in Select Star with the same email, their different accounts will be listed under their [User Profile](/data-discovery/im-new-to-the-team#team-profiles).
  * **Popularity** - Calculated by Select Star. This is based on impressions. For dashboards it is the # of views over a period of 30-90 days.
  * **Loading Status** - This shows whether or not the dashboard has run successfully.
    * :white\_check\_mark: Indicates the dashboard is loading successfully
    * ​:question: Indicates uncertain loading status. This dashboard has not been run in a while
    * :warning: This dashboard's last run was unsuccessful. One or more of the queries has an error.
  * **Created** - When the dashboard was created
  * **Last Run** - When the dashboard was last run. This is currently updated during Select Star's daily metadata sync. If you run your dashboards more than once a day, this may not show the most recent run.
  * **Last Update** - When the metadata of the dashboard was last changed in the BI Tool.
  * **Mentioned By** - Where in Select Star this dashboard was mentioned.
  * **History** - The user activity of this dashboard in Select Star
* **Tiles (**[**Looker**](/learning-data/getting-started-looker)**)/Views (**[**Tableau**](/learning-data/getting-started-tableau)**)/Queries (**[**Mode**](/learning-data/getting-started-mode)**)** - Components of the Explore, Workbook, or Dashboard.
* [**Lineage**](/features/lineage) - Works the same way as the Lineage tab for tables. You can explore Upstream Sources and Downstream Targets. This will have the `Dashboards` option at the bottom of the screen checked by default and centered on the current dashboard.
* **Top Users** - The people or service accounts who query or view this table most.
* [**Discussion**](/features/discussion) - A place for anyone in Select Star to post searchable questions or comments about the dashboard.


# Data Lineage

Understand data lineage in Select Star to explore relationships and evaluate impact. Learn more about lineage in our docs.

Lineage is an important part of understanding your data ecosystem, in this page you will learn how to:

* Understand how different parts of your data relate to each other
* [Evaluate impact of changes](/data-discovery/change-management) to upstream or downstream data sources
* Filter your data lineage by Data Type, Search Term, and by Excluding your Search Term

Watch this video for detailed information on using the lineage feature, or continue reading for an overview. Or skip to the next section to read about Lineage in more detail.

{% embed url="<https://www.loom.com/share/4bd44507d1b842c8a0c13f6d4a20e744?hideEmbedTopBar=true>" %}

## Lineage

Select Star can show you **column-level lineage** for your data assets. The lineage view is designed to show where the data is coming from and where is it flowing towards, so you can find dependencies of each table, column, or dashboard, and see how changes to your assets would impact your data environment.

When you connect a data source to Select Star, Lineage is automatically generated by parsing the SQL statements that ran in your data source.

There are 4 different views of lineage we show out of the box:

1. Upstream: Shows the immediate upstream dependencies in a tree hierarchy
2. Downstream: Shows the immediate downstream dependencies in a tree hierarchy
3. Downstream Dashboards: Shows all dashboard dependencies downstream with extended information like Top User, or dashboard Popularity.
4. Explore: Shows an advanced lineage graph that allows to navigate the flow of data at a column level.

{% hint style="info" %}
There are many ways to see lineage: Check out our [Rest API](/select-star-api), or click on a column from the Column view to start exploring more advanced ways.
{% endhint %}

<figure><img src="/files/5fHLmQ3eqXEywiz9WUaf" alt=""><figcaption></figcaption></figure>

### Lineage Graph

The Lineage graph shows the Upstream Sources and Downstream Targets of the data asset. You can explore the graph by (1) clicking on the tree hierarchy displayed on the left hand side, or (2) by clicking on each of the nodes and columns.

Please note that not all columns are shown in the lineage graph. Select star only shows columns that have any lineage. If a column has no lineage, it is not shown in the lineage graph.

<div data-full-width="true"><figure><img src="/files/4W31uKppbrCw8IneNjkO" alt=""><figcaption></figcaption></figure></div>

### Lineage Search

Search is available within the Lineage Graph too, so exploring tables that have a large number of columns is easier. Follow the instructions below to show the search.

1. Pin any given node by clicking it
2. Click on the magnifying glass icon
3. Type the term you are looking for

The search is available wherever you can find a magnifying glass icon. If you want to do a wider search through the whole graph, you can use the tree hierarchy to search through all the nodes.

<figure><img src="/files/Okx0A7ka6gAdrGxUttvR" alt=""><figcaption></figcaption></figure>

## Filtering Data Lineage

If you need to narrow down results, use some of our filtering features on your upstream or downstream lineage:

* Filter by:
  * Data Type
  * Search Term
  * Exclude Search Term

Note that you can layer the **Data Type** *and* **Search Term/Exclude Search Term,** but Search Term and Exclude Search Term are not able to be layered/applied simultaneously.

Open your 🔍 *Filtering Options* to get started:

![](/files/prnVKaajXSdJpsoEW8fH)

From there, you can **Search by Term**:

![](/files/zx4UoZB87MkE7wDQfm1J)

**Search by Term** and **Filter by Data Type**:

![](/files/RXEcRHMzLAyd4IL34kT1)

And **Exclude Search Term**:

![](/files/5RP4Wilg3yAiymLgjJJR)

## Types of Data propagation

When talking about lineage, we say that data is propagated downwards to downstream data asset (another table, view, dashboard, etc). Data can be propagated as follows

* `AS IS`: The data in the target is identical in value and format to that in the source.
* `AGGREGATED`: The data in the target has been aggregated and the value in target may be different from the one at source.
* `TRANSFORMED`: The data in the target has been aggregated and the format and values might be different from the ones at source.

When calculating lineage between your assets, we also automatically classify downstream propagation. You can see how a column is propagated by editing the column tags. Learn more about tagging in [Tag Management](/data-management/tag-management).

## Lineage FAQs

#### How often does lineage refresh?

Lineage refreshes approximately every 24 hours, after metadata sync is complete.

#### How does Select Star detect updates to lineage?

Select Star looks at DDL statements (used to build and modify the structure of your database) and DML statements (used to query and modify the data in your tables) to identify the lineage of your data.

Select Star will add new relationships to lineage based on both DDL (e.g. CREATE) and DML (e.g. INSERT/UPDATE) statements, however will only remove lineage relationships if a new DDL statement is detected.


# Entity Relationship Diagram (ERD)

Explore Select Star's auto-generated ERD to understand data models and table relationships. View ERDs on table pages or create custom ERDs.

Select Star automatically generates ERDs (Entity-Relationship Diagrams) so you can easily see how tables and columns in your data model are connected, whether through joins or foreign key relationships.

<figure><img src="/files/3SUQH5Mt1etmS9F71b4E" alt=""><figcaption></figcaption></figure>

## How ERDs are calculated

Entity Relationship Diagrams (ERDs) show how tables and columns relate to each other.

In traditional SQL databases, these relationships are usually defined through Primary Key (PK) and Foreign Key (FK) constraints. But many modern data warehouses don’t enforce these constraints, even if they support them. That’s why ERDs in Select Star use more than just metadata.

We determine table and column relationships using two sources:

* **Database metadata:** PK and FK constraints defined in the schema
* **SQL query history:** Actual JOIN clauses from past queries, which show how data is *really* being used—even when no formal constraints exist

This approach highlights not only defined relationships, but also inferred ones based on real usage.

Want more detail? Check out our blog post on [Automated Data Documentation and the Entity-Relationship Diagram](https://www.selectstar.com/blog/automated-data-documentation-entity-relationship-diagram-erd), where you can read more about this approach.

## Table ERD

Every table or view has an associated ERD.

To open it:

1. Go to the **Table** page
2. Click the **ERD** button in the header

<figure><img src="/files/kD32cAu97bXhp3BfcK7U" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
If a table has no metadata information about Primary and Foreign Keys, and there are no relevant joins in the query history, the ERD will appear empty.

That doesn’t mean the table isn’t related, it just means we don’t have enough info to infer any relationships.
{% endhint %}

## Multi-Table ERD

To view relationships across multiple tables:

1. Go to the **Database** or **Schema** view
2. Select the tables you’re interested in using the checkboxes
3. Click the **ERD** button in the header

<figure><img src="/files/2dLdKx0nZsWbcEXF9ITZ" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
Only related tables will appear in the ERD. If there’s no connection between selected tables, they won’t show up in the diagram.
{% endhint %}

## Removing ERD edges

You can remove any edge from the ERD diagram to keep it clean and focused. This is useful when certain joins or relationships are no longer relevant or clutter the view.

To remove a connection from the ERD:

1. Hover over the edge in the diagram
2. Click the edit icon
3. Select **Remove Edge** and confirm the action

Removing an edge will also remove the associated Foreign Key or JOIN relationship. That connection won’t appear in the ERD again — even if it shows up in future query history. Removed joins also won’t appear in the Popular Joins tab or in warehouse query-based suggestions.

<figure><img src="/files/fiXpeN6x5sMbCbyKfIru" alt=""><figcaption></figcaption></figure>


# Queries & Joins

Explore the Queries & Joins tab on tables to understand data context and learn about your team's data.

The Queries & Joins tab appears on database tables to show common queries performed on the table.

You can use the Queries & Joins tab to answer questions like...

* [How can I get the full context of this data?](/data-discovery/how-can-i-use-this-data)
* [I'm new to the team](/data-discovery/im-new-to-the-team) - How can I learn about our data?

![](/files/-MiYzbeIDAsJON2O_jcR)

## Popular Queries

The **Popular Queries** tab shows the queries run most frequently on the table, sorted by popularity.

![](/files/-MgC7c-x6vZS4EPYLYtg)

If you have BI or ETL tools with accounts which regularly run queries, this may skew popularity. [You can mark these accounts](/data-source-management/connect-data-source-users-to-select-star#what-is-that-service-account-checkbox-for) so they will not affect popularity if desired.

Click on the **Query Name /Sample** to see detailed information about the query.

![](/files/-MgC7c-yJO-JdhKHNmaM)

The right sidebar highlights helpful information information at a glance, such as the different tables used in the query and join conditions.

Copy the SQL from the query details and run it in your data warehouse to see what kinds of results it returns.

## Popular Joins

![](/files/-MgC7c-zD7c2h01UtnkR)

The **Popular Joins** tab shows the tables that are most frequently joined to the current table. Click the Join Condition to see an example of a query used to join the current table to other tables.

![](/files/-Mhu01_oIBypxM-FN4uq)

## Recent Queries

![](/files/-MgC7c0-XL_dKNVqYH-U)

The **Recent Queries** tab shows the 50 most recent queries run on that table since the last metadata sync. This can give a broader view what people in your organization use this dataset for.

{% hint style="info" %}
:keyboard: You can filter these results by User or Query using the `Cmd/Ctrl+F` shortcut on your keyboard.
{% endhint %}


# Tags

Use Select Star tags to organize your data by category, status, depricated, sensitive and more. Find the data that matters quickly.

This page has information about how Tags can be used to organize data in Select Star.

* [All Tags](#all-tags)
* [Tag Page](#tag-page)
* [Finding Tagged Data](#finding-tagged-data)
* [Suggested Tags](/features/auto-documentation#suggested-tags)
* [Nested Tags](#nested-tag)

{% hint style="info" %}
Are you an [Admin](/user-management/user-roles) looking for information on how to create and manage tags? See [Tag Management](/data-management/tag-management).
{% endhint %}

## Tags

Admins can create custom Tags in Select Star to organize your organization's data.

Tags can be used to indicate data quality, or any other trait relevant for your organization. Mark any datasets or fields as **Certified**, **Sensitive**, **Deprecated**, or any other status you'd use to define the state of the data.

## All Tags

Find the All **Tags** page in the left sidebar to see the different tags created by your organization.

<figure><img src="/files/fJoL7TBPnDeb2Nkcy08u" alt=""><figcaption></figcaption></figure>

For example, the Certified tag contains datasets have been created, approved, and vetted by data experts in your organization. This is a great place to start for [newcomers](/data-discovery/im-new-to-the-team) to start and looking to get familiar with your data.

<figure><img src="/files/Nsx6UySjfPBwAz7ghV1W" alt=""><figcaption></figcaption></figure>

Each Tag can have sub-tags associated with it. You can see the hierarchy of tags in the sidebar or on the that particular tag page

<figure><img src="/files/vJCJUi49XMvk7KWCeOz7" alt=""><figcaption></figcaption></figure>

## Tag Page

Each Tag Page can have multiple different types of assets associated with it. To get an overview of all the different assets assigned to a Tag, you can go to the assets tab:

<figure><img src="/files/WM3gsXhBdh9a7meE6sMW" alt=""><figcaption></figcaption></figure>

You can also ask questions in the discussion tab for each tag to ensure all related information close by.

<figure><img src="/files/MZACckAVEQlzVNKnTdEf" alt=""><figcaption></figcaption></figure>

## Finding Tagged Data

### Search

You can see all tagged data another way by searching for the tag in the search bar.

![](/files/-MgC87vLctmcEgMWtYYH)

### Filter

Filter data with specific tags in [database](/data-discovery/wheres-my-data) pages or [dashboard](/data-discovery/wheres-my-dashboard) pages as well.

<figure><img src="/files/g5jTKGv9nK1Tvjeu8irv" alt=""><figcaption></figcaption></figure>

### Nested Tags

In order to nest a tag, just drag it and drop it over the tag you wish to nest it under.


# Teams

"All Teams" page lists organization's teams. Team pages show details, members, favorites, assets, and discussions. Search and mention teams.

Teams are used to organize your Select Star members and see usage across different cohorts of users. Teams can also be used for defining access policies within Select Star with [Policy Based Access Control](/user-management/policy-based-access-control).

## Video Tutorial

{% embed url="<https://www.loom.com/share/80422d625b0b4bc1a01a24b0f983f9b9?hideEmbedTopBar=true>" %}

## All Teams

On this page you will see a list of all teams created by your organization, and once you click on a specific team name, you will navigate to that team's page.

<figure><img src="/files/qCZmoHaDt5K0GfD8ZxRk" alt=""><figcaption></figcaption></figure>

## Team Page

On this page you will find all the details about that team and their members.

<figure><img src="/files/UMV83Dw0tASE6XcIGdOW" alt=""><figcaption></figcaption></figure>

### History

This section shows all the activity on the Select Star platform, for all the team members.

![](/files/J8ZAelvOAMsxdRtIz2XX)

### Team Description

This sections allows you to add your team description, mention resources, like team Slack channels or GitHub repo links, and even related data assets, like tables or dashboards. This will be visible on the All Teams Page.\\

<figure><img src="/files/Pzq6y3NVj766bGgJN1XJ" alt=""><figcaption></figcaption></figure>

### Team Members

On the Team Members tab, you will see a list of users that have been added to this team by an administrator.

Using the search feature, you can search for a team member by name or email id. You can also use the sorting buttons to organize the table in a way that's more convenient.

The role refers to a specific function and responsibility in Select Star that is assigned to each team member. To learn roles visit this [page](https://docs.selectstar.com/user-management/user-roles).

### Favorites

On the Favorites tab, you can see what data assets that team's members have marked as favorites.

<figure><img src="/files/rtW2RBTx1zVhNoj4qsgG" alt=""><figcaption></figcaption></figure>

### Owns

On the Owns tab, you can see all the data assets owned by the team's members, including, Tables, Dashboards, Metrics and Tags.

<figure><img src="/files/7zammrwuiEnRbK42WxfE" alt=""><figcaption></figcaption></figure>

### Discussion

On the Discussion tab, you can leave a comment for that team's members.

<figure><img src="/files/l954hsaVOPfsO2eZiMgn" alt=""><figcaption></figcaption></figure>

## Finding Teams

### Search

Using the top search feature, you can search and navigate to the team within Select Star.

<figure><img src="/files/WfEHNrdT6wiD2deAMyx2" alt=""><figcaption></figcaption></figure>

### Mentions

To use mentions, you can simply type `@` and the team's name, in order to search and mention them.

<figure><img src="/files/597uXbLVxcWmfulWwYiD" alt=""><figcaption></figcaption></figure>

The Mentioned by section in the Team's page provides you with the information about which asset the team was mention in.

<figure><img src="/files/OGE46XQqrB3bR2LL7qb3" alt=""><figcaption></figcaption></figure>

## Team Notifications

Like users, teams can be assigned owners of objects and mentioned in descriptions and comments. By default teams will receive notifications for:

* Metadata changes
  * You are added as an [Owner](/data-management/data-ownership) to a data asset
  * Someone [comments](/features/discussion) on a data asset you own
  * Someone modifies the description or tags on a data asset you own
* Schema changes
  * There is a schema change in an object that you own
* Mentions
  * Someone replies to a comment you posted, or mentions you in a comment

Admins can modify team notifications settings in **Admin** > **Teams.** Usually all members of a team will receive notifications for the team. If an email or slack channel is set to receive notifications for the team, individual team members will no longer be notified directly.

Note: The Select Star slackbot must be invited to a channel to be able to send notifications for that channel.

<figure><img src="/files/wii0bFAgUtVZU3bDlB1n" alt=""><figcaption><p>Modify team notification settings</p></figcaption></figure>


# Discussion

Learn how to utilize Select Star's Discussion tab for efficient dataset management, knowledge sharing, and cross-team collaboration.

You can use the Discussion tab to...

* [Manage changes](/data-discovery/change-management) to datasets
* Provide a place for new employees to ask questions and learn about data
* Allow people from different teams to ask direct questions about unfamiliar datasets and get responses from experts

## Video Tutorial

{% embed url="<https://www.loom.com/share/8c3737f9074f4701b4b83e2408bdf892?hideEmbedTopBar=true>" %}

## Creating Discussion Items

Click the **Discussion** tab from any [dashboard](/features/dashboard-page) or [table page](/features/table-page).

![](/files/-MiYnz9r6498fK52h5JG)

Use `@` to bring up a search in your comment and notify users, or even create clickable links to other Select Star data assets.

![](/files/H6tW0i0FHKrTS0jtNemw)

![](/files/jeevumzD09dASmZKJJhK)

The [Business and Technical Owners](/data-management/data-ownership) of the table or dashboard will be automatically notified when a comment is created. Notification emails are sent by default, but can be disabled in Settings.

Notifications are also sent through Select Star. See notifications by clicking the bell :bell: icon in the top right corner of the screen.

You can also receive notifications by installing the [Select Star Slack App](/integrations/slack).

![](/files/2xTdCURDvy78Kmk6kUWa)

## Discussion Tips

Table Owners can pin :pushpin: important comments so they'll show up at the top of the page.

![](/files/-MgvAGVN104HveLwWEa3)

Discussion comments show up in search results under the **Select Star** tab.

![](/files/Nez4QzHWdRRX9F7dkkGj)


# Downstream Notifications

Communicate key updates and upcoming changes to downstream users and key stakeholders easily with downstream notifications.

Select Start supports downstream notifications to be able to quickly reach out to the top users of data assets or downstream owners when there is an update or upcoming breaking change.

Downstream notifications are supported for Database Tables and BI Dashboards. The following groups are supported:

* Database Tables
  * Downstream Owners - This is anyone who owns a table (including BI models) or dashboard
  * Table Users - Top 25 users with the most queries against the table (including data source users who may not be linked to a Select Star user)
  * Dashboard Users - Top 25 users with the most views/queries of dashboard directly downstream from the database table
* BI Dashboards
  * Top Users - Top 25 users with the most views/queries of dashboard directly downstream from the database table

Note: the users in these groups can also be seen in the coresponding tab in **Top Users** for the object.

Users must be an admin or owner of the object to send a downstream notification.

## Creating a Downstream Notification

Create a downstream notification from anywhere on the object with the **Notify** button on the top right. This will automatically take you to the **Discussion** tab and open the notification modal.

<figure><img src="/files/ap6WhXePMP7SmEjnAEkP" alt=""><figcaption></figcaption></figure>

Users can also navigate to the **Discussion** tab and create a notification when drafting your message.

<figure><img src="/files/uUrHNAYYdtej5RXLcIhu" alt=""><figcaption></figcaption></figure>

Users can select a recipient for the notification, choosing from one of the preset groups, or adding additional users and teams to the list.

<figure><img src="/files/rb3NdoPRAvZ3UiofCc3T" alt=""><figcaption></figcaption></figure>

The author can see which users are in a group by hovering over that group or in the **Top Users** tab on that object.

After selecting recipients and crafting your message, click **Send** and you're done. Users will receive downstream notifications in-app as well as in their email (if they have email notifications enabled).


# Documentation

Leverage Select Star's Documents feature for comprehensive data documentation. Use rich text editing, nesting, and mentions for effective knowledge organization.

Select Star has three types of documentation:

1. **Pages** - Pages are general purpose and can be used from anything from long technical documentation about processes and data models, to describing key business details at great length, or even an onboarding or feature request doc.
2. **Glossary** - The glossary is the place to document key business terms that are common across your organization. The glossary ensures that everyone is using the same language to facilitate effective communication. e.g. ARR, Subscription, MQL
3. **Metrics** - Metrics are a defined calculation that has business meaning or aligns with a business goal. e.g. ARR could be a metric as well as a glossary term - the glossary term defines the business meaning of the term, and the metric defines how it is calculated and where in the data stack it is represented

### Formatting

All document types generally allow for rich text editing, including **headers, numbered lists,** and **code blocks**. You can **mention** other documents, metrics, users, tables, columns, and BI elements.

Use @ symbol across Select Star in documents, descriptions, and discussions to reference other objects, to create a knowledge network you can navigate. \\

<figure><img src="/files/Ckc1lqp7KBfGxS2PFAS6" alt=""><figcaption></figcaption></figure>

If the mentioned object is deprecated, we will show the "Invalid mention" message rather than the general mention format. We do not keep deleted data object names in mentions, but the names of any document type object (page, metric, or glossary) will remain even after deletion.\\

<figure><img src="/files/XcKyifnpajSlEp48Ukuv" alt=""><figcaption><p>Invalid mention examples</p></figcaption></figure>

### Document Organization

Nest documents in the hierarchy in order to organize them. Drag it and drop it over the document you wish to nest it under.

<figure><img src="/files/3wbLkAwjoXK7o01cBBDT" alt=""><figcaption></figcaption></figure>


# Pages

Leverage Select Star's Documents feature for comprehensive data documentation. Use rich text editing, nesting, and mentions for effective knowledge organization.

Use Select Star to create and maintain your data documentation using our Documents feature.

![](/files/nsJQVWkXMuVM05zMlBoK)

Pages will not be editable until you explicitly click the **Edit** button.

![](/files/1KZKrBb2X92Mw6RPgls4)

Pages allow for rich text editing, including **headers, numbered lists,** and **code blocks**. You can **mention** other documents, metrics, users, tables, columns, and BI elements.


# Metrics

Select Star offers a range of metrics that are based on data and are designed to answer specific business questions. Learn how these metrics can help you make informed decisions and drive success.

**Metrics** are calculations based on data that aim to answer specific business questions.

Customize the columns using the ![](/files/2qSgPtoqlNo1YzsBK7Tw) icon in the header, or use the keyboard shortcut `Cmd/Ctrl+F` to open a search filter.

<figure><img src="/files/0TqcuqZyRgydieLwT2Cm" alt=""><figcaption></figcaption></figure>

To add a new metric, click **+ New Doc -> Metric**.

Name your metric.

You can create an empty metric, or search for a measure field that represents the metric.

<figure><img src="/files/4KizykDFYcse8txE2Ari" alt="" width="375"><figcaption></figcaption></figure>

{% hint style="info" %}
Note: When you choose a Column/Fields to Represent a Metric, this Column will now display with a 'Metric' label:
{% endhint %}

<figure><img src="/files/SqyPfj0GanGurkESoDE3" alt=""><figcaption></figcaption></figure>

Then click the 'Metric' label to see everywhere this Column represents a Metric:\\

<figure><img src="/files/mq6qxVRyaKx4p4T6vrbj" alt="" width="563"><figcaption></figcaption></figure>

Select Star can automatically detect dimensions that you may want to group or pivot on, but you can remove these if desired or search for others.

<figure><img src="/files/nJvQBaSd6Go5kYs494Mk" alt="" width="375"><figcaption></figcaption></figure>

Save the metric to see new fields where you can provide a **Description**, **Business Questions to Answer**, and **How it's Calculated** as editable fields.

<figure><img src="/files/63Q6wcDV2Tplen0ec2jy" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
Type `@` and then a few characters to create a clickable link to a specific data asset within your calculation or description.
{% endhint %}

The **How it's calculated** field is meant to give a high-level overview of the metric to business-oriented stakeholders.

If you have a SQL query which will show a more accurate representation to technical stakeholders, you can add it where you'd modify the [Owners](/data-management/data-ownership) by clicking **< > SQL**.

<figure><img src="/files/v8Q3Kz6XvfVDHBg7w48A" alt=""><figcaption></figcaption></figure>

You'll be able to edit or paste in a SQL query.

![](/files/-MiOKtnf2IUD5tUbF1Ep)

Stakeholders can also use the [Discussion](/features/discussion) tab to ask questions or make comments about the metric.

The owners will be notified of any comments. By default, both the Business and Technical Owner of a metric are set to the person who created it. [Data Managers and Admins](/user-management/user-roles) can change the Owners as needed, and even apply [Tags](/features/tags) to metrics by clicking the 🏷 icon next to the title.


# Metrics Generation

{% hint style="info" %}
Metrics Generation is currently a Beta feature. If you have any questions or feedback, please reach out to <support@getselectstar.com>
{% endhint %}

Select Star can help you document metrics that currently exist in your BI layer, saving hours spent documenting.

By looking at the calculation and lineage behind your dashboards, Select Star can generate metric documentation, including the formula, dimensions, and where it's represented in your BI layer. We extract the measures in the datasets powering your dashboards

{% hint style="info" %}
Metrics generation is currently supported for PowerBI and Tableau with support for Looker coming soon.
{% endhint %}

<figure><img src="/files/Yin9LGwveZrF3hKWgDFq" alt=""><figcaption></figcaption></figure>

## Generating Metrics

The process for generating metrics is as follows:

1. On the dashboard list page, select one or more dashboards that you want to generate metrics for
2. Click **Generate Metrics** at the top of the list

<figure><img src="/files/SC1Jyi1heNp75DD8Ob6W" alt=""><figcaption></figcaption></figure>

3. Click confirm in the modal. Metrics will be created with the **Draft** tag applied, and you will receive an email summary with metrics that were created.

<figure><img src="/files/6AmZawwbS3bx3e3BPKVp" alt=""><figcaption></figcaption></figure>

4. Review the generated metrics.

Once metrics are finished generating, you'll receive an email with the generated metrics. We highly recommend you review the generated metrics and organize them into [Collections](/data-management/collections) and remove the **Draft** tag when ready.

<figure><img src="/files/dIOsrt0eLOUn4wbYzZKw" alt=""><figcaption></figcaption></figure>

### Things to Know

1. The user who generates the metric will automatically be set as the business and technical owner of the created metrics
2. If Select Star detects that a metric already exists, it will not create a duplicate metric, and skipped metrics will be indicated in the email. Duplicates are detected when:
   1. Another metric with the same data measure defined in it's "represented as" section already exists
   2. If another metric already exists with the same name
   3. If mulitiple generated metrics share the same name, only one will be created, with all measures included in the "represented as" field
3. If Select Star detects a similar metric, it will be included in the generated metric's description


# Glossary

The **glossary** helps you keep key business **terms** centralized and common across the organization.

All users can view Glossary terms in the Docs tab, and additionally can filter to Glossary terms under All Documents.

<figure><img src="/files/SCPl3weyi126hIaKceF3" alt=""><figcaption></figcaption></figure>

Data Managers and Admins can see and create new glossary terms directly in the UI by clicking on **Create New Doc > Glossary** in the left navigation, or can upload a CSV to create and update terms in bulk.

<figure><img src="/files/ny5XKmXH7qXoIPip9JqA" alt=""><figcaption></figcaption></figure>

When creating and updating Glossary terms via CSV upload, note the expected format is

* Columns in the CSV:
  * **parent\_guid** - The guid of the parent document for this glossary term, which is used to set the hierarchy of documents. If empty, the glossary term will be created at the top level with no nesting.
  * **guid** - The guid of the glossary term. If no guid is given in an import, a new term will be created. If the guid is present, that term will be updated.
  * **name** - The name of the glossary term.
  * **description** - The description of the glossary term, in markdown format. Note: this does not support @ mentions format within Select Star.
  * **business\_owner** - The business owner of the term. For users this is the user email, for teams this is the team name.If no owner is given when creating a new term, the user performing the upload will be set as the owner.
  * **technical\_owner** - The technical owner of the term. For users this is the user email, for teams this is the team name.If no owner is given when creating a new term, the user performing the upload will be set as the owner.
  * **tags** - The tags associated with the glossary term, as a comma separated list.

Stakeholders can also use the [Discussion](/features/discussion) tab to ask questions or make comments about the glossary term.

The owners will be notified of any comments. By default, both the Business and Technical Owner of a glossary term are set to the person who created it. [Data Managers and Admins](/user-management/user-roles) can change the Owners as needed, and even apply [Tags](/features/tags) by clicking the 🏷 icon next to the title.

Useful links:

* Glossary terms API documentation: <https://api.production.selectstar.com/docs/#tag/terms>


# Automated Documentation

Simplify documentation with Select Star's automated features: suggested descriptions, AI-generated documentation, suggested tags, and public documentation loading.

Select Star makes documentation simple with suggested documentation, description and tag propagation, and public documentation automatically loaded.

Suggested description and suggested documentation are partially based on how data is propagated downstream. You can learn more about how we classify downstream propagation in the [Data Lineage - Types of Data Propagation](/features/lineage) section.

{% hint style="info" %}
We base suggestion on how data is documented downstream, but we tag in both up and downstream. We suggest you add descriptions and tags at the source and let the propagation be downstream, but we will also do the process upstream.
{% endhint %}

## Suggested Descriptions

Select Star will automatically fill in table & column descriptions that match the following:

* Upstream / downstream fields where the data has been inherited "**As Is"**
* Similar columns (table structure is the same and we regard the table as a duplicate)
* Similar BI fields

<figure><img src="/files/K5WzfT1oeYqjzmTe4Puf" alt=""><figcaption><p>Suggested and user descriptions for table columns</p></figcaption></figure>

User descriptions and loaded descriptions will be displayed in black, and suggested descriptions will be shown in gray to highlight the origin of the description.

You can edit the description to set its value manually and this will override the suggested description.

## AI Generated Documentation

Select Star can automatically generate table, column, dashboard and chart descriptions using Open AI's and Anthropic's GPT models, based on the metadata, SQL queries, existing documentation, dashboard and chart titles and formulas, and other data context that Select Star has analyzed.

Data managers can generate descriptions directly on an object by clicking on the icon beside the description, which will kick off the AI generation. Once the AI description is generated, users can select it in the modal to fill in blank or replace existing descriptions.

<figure><img src="/files/SraRqxUhBPcNtYV7ExGO" alt=""><figcaption><p>AI generated table description</p></figcaption></figure>

You can also generate AI descriptions in bulk. With AI Assist, you can fill in all columns in a table without descriptions with a single button click. You can also remove all generated description in one click.

<figure><img src="/files/hkrRuYPmAssyxSr6repX" alt=""><figcaption><p>Bulk generate and bulk remove AI descriptions</p></figcaption></figure>

## Suggested Tags

Tagging is one of the most important feature in data cataloguing and data governance. Select Star will automatically suggest tags in your data assets so you don't have to tag every single element in your data catalog.

We infer tags for suggestions by

* Upstream and downstream fields where data is propagated `AS IS`
* Tables and column relationships for PII.

<figure><img src="/files/UlRNNznEyPLKEn3OFo2b" alt=""><figcaption><p>Suggested tags - auto filled from another Table</p></figcaption></figure>

Suggested tags will show up with a grey background to indicate they are suggested and not manually assigned by a user.

You can hover over the tag to get more information about the origin of this suggestion.

Tags can be edited at any time so you can explicitly override tag suggestions to suit your needs.

## Public Docs

When using tables where public documentation is available, Select Star automatically loads that documentation into your Select Star instance.

![](/files/D2DrmhIfJnmnPG4iN2yH)

Descriptions loaded from public documentation will display as gray text until a user makes an edit.


# User Analytics

With the Select Star Analytics page, organization administrators can easily track user and page metrics to gain valuable insights into user behavior and interaction with the platform. Discover more ab

*Want to see how popular your tables, columns, dashboards, and other assets are with your team? Head over to* [*Popularity*](/data-discovery/how-can-i-use-this-data#popularity)*.*

## **Introduction**

Under the Analytics page, organization admins will be able to track user and page metrics. All statistics are based on user activities that are logged when they perform an action within Select Star.

To access the Analytics page, click **Setting** from the menu on the top right, Analytics section is available under the **Admin** section on the left sidebar.

<figure><img src="/files/p4Tr0VHT8EqDn3ejYO87" alt=""><figcaption></figcaption></figure>

## Date range filtering

Every analytics page has a date range picker to filter metrics by specific dates. For now maximum date range limit is 90 days. The default range is set to last 30 days.

To change the dates, click on the dates range, and select the first and the last date and hit **Confirm**.

<figure><img src="/files/7eqvNyPy8G0L4vv7Hu1C" alt=""><figcaption><p>Date range filter</p></figcaption></figure>

We also have a predefined date range to select from. The selected range will be displayed at the bottom of the modal. If you want to apply that date range just click confirm button.

<figure><img src="/files/cwGRa9TSsKJ4hCfVb5aH" alt=""><figcaption><p>Predefined Date Range</p></figcaption></figure>

## Analytics - Users

*Users analytics lets you track metrics about your organization's users within Select Star.*

<figure><img src="/files/3GwyeYx1YmNB2nD82Les" alt=""><figcaption><p>Analytics - Users tab</p></figcaption></figure>

#### Overview Metrics:

* **New users**: How many users were created during that time period.
* **Active users**: How many users performed an action during that time period.
* **Total Members**: How many users were registered on the system up to the last day of the date range

### **Top Active Users**

Top 5 Users who performed the most actions, click **Show more** to see the full list and **Activity Count**.

<figure><img src="/files/Kgrp2NedhovYrA5fyhGw" alt=""><figcaption><p>Top 5 Active Users</p></figcaption></figure>

### **Top Contributors**

Top 5 Users who performed the most actions, excluding views, click **Show more** to see the full list and **Edit Count**.

<figure><img src="/files/4JP4Q6nA5eZEa19eZkKL" alt=""><figcaption><p>Top 5 Contributors</p></figcaption></figure>

### **New Users**

Top 5 New Users based on Join Date, click **Show more** to see the full list.

<figure><img src="/files/D49AkKqJR5yiWsP8GuS5" alt=""><figcaption><p>Top 5 New Users</p></figcaption></figure>

## Analytics - Activities

*Activities analytics lets you track metrics about your organization's pages and search terms within Select Star.*

<figure><img src="/files/dJR9hyB9wc6ywVaNTR23" alt=""><figcaption><p>Analytics - Activities tab</p></figcaption></figure>

#### Published Resources Metrics:

* **Table Description Fill Rate**: Percentage of total tables for which a description was filled in.
* **Column Description Fill Rate**: Percentage of total columns for which a description was filled in.
* **Total Number of Docs**: Total count of documents within Select Star.
* **Tables with Owners Rate:** Percentage of total tables for which ownership (technical or business) was filled in.
* **Total Number of Discussions**: Total count of discussions (does not includes replies) within Select Star.

{% hint style="info" %}
Hover over the tooltip to read how the metrics were calculated.
{% endhint %}

### **Top Pages**

Top 5 viewed pages, click **Show more** to see the full list.

<figure><img src="/files/8tldZy7UKJqucjaz8bGI" alt=""><figcaption><p>Top 5 pages</p></figcaption></figure>

### **Top Pages** **Edited**

Top 5 edited tables, columns, dashboards and charts within Select Star, click **Show more** to see the full list.

<figure><img src="/files/BUY6JpWaTcpMHRj62mP9" alt=""><figcaption><p>Top 5 pages edited</p></figcaption></figure>

### **Top Search Terms**

Top 5 used keywords for searching, click **Show more** to see the full list.

<figure><img src="/files/SOKtL4jBw8nLCk0gz9ip" alt=""><figcaption><p>Top 5 search terms</p></figcaption></figure>

## Analytics - Activity Log

*Activity log lets you track all user's actions within Select Star.*

<figure><img src="/files/RznCsTZZ8s6ZVUskyoTe" alt=""><figcaption><p>Analytics - Activity Log tab</p></figcaption></figure>

### Export to CSV

Export to CSV button will let you download with all log data for the selected date range.

<figure><img src="/files/dNIhejT1nxOYMUfIkIht" alt=""><figcaption></figcaption></figure>


# Chrome Extension

The Select Star Chrome Extension allows users to get their data context side-by-side their dashboards. It familiar web interface with the added benefit of one click search.

The Select Star Chrome Extension allows users to get their data context side-by-side their dashboards and tables. With the Extension, users have easy access to their metadata while they're looking at dashboards. This allows users to keep their focus in one place, removing the need to switch between tabs or applications.

Users can install the extension individually by following the steps below, or admins can perform an [organization-wide install](/features/chrome-extension/organization-wide-install) for all users.

## Video Tutorial

{% embed url="<https://www.loom.com/share/84556c72d48e43f2957f2a9761108e14>" %}

## How to install?

**Step 1.** Go to [the Select Star Chrome Extension web store page](https://chrome.google.com/webstore/detail/select-star/okjjmjmiaemccchgkpchjpkaomabndek)

<figure><img src="/files/dQk3bTGYsrUJTkhsclOV" alt=""><figcaption><p>Chrome web store page</p></figcaption></figure>

**Step 2**. Click **Add to Chrome**

<figure><img src="/files/Um1aB2dROIuqrnHTSZXK" alt=""><figcaption></figcaption></figure>

And that's it. The extension should be available to use after you log in via the extension.

## Main Features

### **One-Click Search**

<figure><img src="/files/aadyaTf7cgnlOjRfbZGf" alt=""><figcaption></figcaption></figure>

### **Search**

<figure><img src="/files/QUcErr6La1x9QZtBWgz9" alt=""><figcaption></figcaption></figure>

### **View Information for Tables and Dashboards**

<figure><img src="/files/tQ3Wjo2QTUm7jciflHkU" alt=""><figcaption></figcaption></figure>

### Open Dashboards and Tables with One-Click

Supported tools:

* BI Tools: Looker, Tableau, Power BI, Sigma, Mode, ThoughtSpot
* Warehouses: Snowflake, Databricks

When using BI Tools, you can open dashboards with one click. Select Star detects when there is a matching page in the catalog and will open to that dashboard or table's catalog page.

<figure><img src="/files/JUi7LtL8Zpyw11ttc7xV" alt=""><figcaption></figcaption></figure>

## Unsupported Pages

Some pages in the application are not yet supported. If you try to navigate to any of these pages in the Chrome Extension, it will open up the page in the web app.

Excluded pages are:

* Documentation (metrics, glossary terms, pages)
* Database and schema pages
* BI folder pages
* Collection pages
* Tag pages
* Team and user profile pages
* All settings and admin pages

Note, some table and dashboard tabs are also excluded from the Extension

* Tables: Queries & Joins, Related, Data Quality, Preview, Lineage>downstream dashboards
* Dashboards: Top Users, Data Quality, Lineage > Source Tables


# Organization-wide install

This guide explains how Google Workspace or Microsoft enterprise administrators can deploy the extension automatically to all users in their organization.

## Chrome Extension Rollout (Google Workspace)

{% embed url="<https://www.loom.com/share/0b64125ab32e4fc08cde01fd08af9c72?sid=3697d5de-c8ad-4047-9560-e477fbceff59>" %}
Walkthrough: Setting up forced install for Select Star in the Google Admin Console
{% endembed %}

To deploy the extension to all managed Chrome users in your domain:

#### Prerequisites

You’ll need the following:

* Administrator access to Google Workspace
* Select Star Chrome Extension ID: `okjjmjmiaemccchgkpchjpkaomabndek`

#### Configure extension roll-out in Google Workspace

1. Go to <https://admin.google.com/ac/chrome/apps/user> and log in as administrator.
2. *Optional: Select the Organizational Unit you want to deploy the extension to.*
3. Click :heavy\_plus\_sign: in the bottom right and choose **Add from Chrome Web Store**.
4. Search for the extension by ID **`okjjmjmiaemccchgkpchjpkaomabndek`** and click **Select**.
5. Set the Installation policy to **Force install + pin to browser toolbar**.
6. Click **Save**.

{% hint style="success" %}
The extension will now automatically appear in Chrome for all users in the selected org unit.
{% endhint %}

## Microsoft Edge Rollout (Group Policy)

To deploy the Select Star extension across all managed Microsoft Edge (Chromium) browsers in your organization, use the **Group Policy Management Console (GPMC)** on a domain controller.

{% hint style="info" %}
This method assumes that Microsoft Edge is already set up as a managed browser for your users. If you haven’t configured Edge management yet, follow this guide to get started: [Manage Microsoft Edge extensions in the enterprise](https://learn.microsoft.com/en-us/deployedge/microsoft-edge-manage-extensions).
{% endhint %}

#### Prerequisites for Edge Extension Deployment via Group Policy

1. Your organization uses **Microsoft Edge (Chromium)**.
2. All devices are **joined to your Active Directory domain**.
3. You have access to the **Group Policy Management Console (GPMC)** on a domain controller.
4. You have permission to create or edit **Group Policy Objects (GPOs).**
5. Microsoft Edge is configured as a **managed browser**.
   1. To verify, open `edge://policy` in Edge and check for active policies.

#### Steps to Configure Group Policy (GPO)

1. Open **Group Policy Management Console (GPMC)** on your domain controller.
2. In the left panel, right-click your domain or a specific Organizational Unit (OU), then select: **Create a GPO in this domain, and Link it here…**

   <figure><img src="/files/AzOG81ubFm1EfIInofKQ" alt=""><figcaption><p>Create a new GPO</p></figcaption></figure>
3. Name the GPO (e.g., *Select Star Extension*) and click OK.
4. Right-click the newly created GPO and select Edit.
5. In the Group Policy Editor, navigate to: **Computer Configuration** → **Policies** → **Administrative Templates** → **Microsoft Edge** → **Extensions**
6. Double-click **Control which extensions are installed silently**.
7. Select **Enabled** and click **Show.**
8. Add the following entry:

   ```
   okjjmjmiaemccchgkpchjpkaomabndek;https://clients2.google.com/service/update2/crx
   ```
9. Click **OK** and close the editor.
10. To apply the policy immediately on a device, run: `gpupdate /force`. Otherwise, the policy will apply automatically during the next policy refresh cycle.

{% hint style="success" %}
The extension will now automatically appear in Microsoft Edge for all users in the selected org unit.
{% endhint %}

{% hint style="info" %}
**Additional Resources**

* [Microsoft Edge extension management guide](https://learn.microsoft.com/en-us/deployedge/microsoft-edge-manage-extensions)
* [ExtensionInstallForcelist policy documentation](https://learn.microsoft.com/en-us/deployedge/microsoft-edge-policies#extensioninstallforcelist)
  {% endhint %}

## Need help?

Contact our support team at <support@selectstar.com> if you have questions or need help with deployment.


# Source Tables

Learn about Select Star's source tables and source columns feature, providing column-level lineage to track data origins and flow.

This page has information about how source tables and source columns feature works in Select Star. Select Star can show you **column-level lineage** for your data assets. The lineage view is designed to show where the data is coming from and where is it flowing towards. Both **source tables** and **source columns** are a feature generated based on our **column-level lineage** and show you from which tables or columns the data is coming from directly.

The main distinguishing factor between source tables or source columns feature and upstream lineage feature is that source tables / columns show you the tables from which the data is coming from directly, instead of listing all tables / columns in the upstream. It can be easily described as first tables / columns found in each path when traversing the upstream lineage.

## Source Tables

Source tables can be used for most elements in Select Star that are BI tools elements like Dashboard, Reports, Tableau Tables, Tableau Data Sources, etc. It is generated based on Table level lineage.

Source Tables feature can be accessed when using **Lineage** Tab as showed below:

![](/files/u7TLAn4BOfRNoxZhpSsy)

For given Dashboard, you can see the source tables for the dashboard. Given the image above with Source Tables for a given Dashboard, it was generated from lineage shown on the image below, where you can see that the Source Tables have upstream tables themselves, but we only showed the direct DWH Tables that are feeding data to the Dashboard.

![](/files/oaUOUC28toP65O5NNsjs)

## Source Columns

Source columns feature is generated based on **column level lineage** and can be accessed by clicking any ReportQuery or any other element from BI tools.

![](/files/ksPvGvDQxf7sh4Be96hp)

![](/files/Oo1CHHlKiL1K4eajcMtt)

This feature is similar as Source Tables, but instead of showing the tables, it shows the columns from which the data is coming from directly.


# Cost Analysis

{% hint style="info" %}
Cost Analysis is only available on certain plans. Please reach out to <support@getselectstar.com> to enable or trial the feature.
{% endhint %}

{% hint style="info" %}
This feature is currently only supported for Snowflake (DWH), Sigma and Mode (BI tools)
{% endhint %}

### Introduction

With Cost Analysis, you will gain insights about your Snowflake Spend at the query, dashboard, warehouse and user level. You will be able to view Cost (in Credits) of your big ticket dashboards and identify which queries are the most expensive by sorting them by cost. You will also be able to identify Snowflake Spend for each active warehouse, and have the ability to track Snowflake Spend by teams and users.

To access the Cost Analysis page, click **Settings** from the menu on the top right, and click Snowflake Cost Analysis in the **Admin** section on the left sidebar.

<figure><img src="/files/jJ678HDr0r5ddzowBEIW" alt=""><figcaption></figcaption></figure>

This feature contains 3 tabs

* Overview - This tab provides a high-level summary of total expenditure, cost trends, and key cost drivers to quickly assess the overall cost landscape.
* Data - This tab provides detailed insights on query cost, dashboard cost, and warehouse cost, allowing admins to analyze and optimize resource allocation for efficient cost management.
* User - This tab provides a comprehensive breakdown of the cost by team and user, allowing admins to track and analyze cost allocation across different groups.

### Date Range Filtering

Each Cost Analysis page has a date range picker to filter metrics by specific dates. Currently, the maximum date range limit is 14 days.

{% hint style="info" %}
Dates and Times shown are in UTC timezone
{% endhint %}

### Benefits of Cost Analysis

Our Cost Analysis feature provides a centralized view of your Snowflake compute costs, and enables your organization to gain insights into your usage patterns and identify cost optimization opportunities. This feature provides:

* **Near Real-time Cost Monitoring**: This feature displays up-to-date information on credit consumption, allowing you to track costs in near real-time and react promptly to any unexpected spikes.
* **Cost Breakdown**: It provides a detailed breakdown of costs by warehouses, user, and dashboards, enabling you to pinpoint areas of high expenditure and optimize resource allocation.
* **Forecasting and Budgeting**: By analyzing historical data and usage patterns, this feature can help forecast future costs, allowing you to plan budgets effectively.
* **Query and Dashboard Usage Analysis**: It can help correlate query and dashboard usage metrics with cost, helping you identify expensive queries that can be optimized or tuned for better efficiency.


# Schema Change Detection

Easily see schema changes to your tables and columns and understand how your data is evolving over time.

Schema Change Detection helps you identify when your data has changed, and how it has done so. This can help catch unplanned changes, or let you know about new data you can use for analysis.

## Detected Events

Schema change detection currently surfaces the following changes:

* Table create / delete
* Column create / delete
* Column data type change

{% hint style="info" %}
Note: this information is processed from the schema metadata itself. If an object is renamed, this will be processed as a delete event with the old name and a create event with the new name.
{% endhint %}

## Schema Change Activities

All users can see schema change events directly on the tables themselves. Navigate to the **Overview** tab on a table. On the right hand side under History, users will be able to see the detected events for the table and for columns in that table.

<figure><img src="/files/kdKztmSayqhShPEz4xq6" alt=""><figcaption><p>Column schema changes on a table</p></figcaption></figure>

Admins will be able to see *all* detected schema change activities on the Admin page under **Analytics > Schema Changes.**

<figure><img src="/files/7J0YWdxhR0m968m57SkQ" alt=""><figcaption><p>Search for specific data assets across all schema changes</p></figcaption></figure>

## Notifications

Data asset owners (users and teams) will receive notifications when an asset the own has a detected schema change.

* In-app notifications will show all changes
* Email notifications will be sent per data source after ingestion (if configured)
* Slack notifications will be sent per data source after ingestion (if configured)

<figure><img src="/files/wrYTGLNsgEH3fC7VWKpk" alt=""><figcaption><p>Email notification for schema changes on Snowflake</p></figcaption></figure>

{% hint style="info" %}
Schema change notifications can be configured at the user, team or organization level. See [Slack Notifications](/integrations/slack#slack-notifications) for more details.
{% endhint %}


# Ask AI Chatbot

Select Star's AI Chatbot will answer your data questions with natural language, combing through your metadata so you don't have to.

## AI Chatbot

The Select Star AI Chatbot is extremely useful in providing in-depth and nuanced answers about your data. It can help find the right data assets to work with or get you started on your query. Some top use cases include:

1. Learn more about your data
   1. Answer questions from existing documentation within Select Star, including descriptions, glossary terms and docs.
   2. Answer questions about how certain things are calculated i.e. How is ARR calculated?
2. Find the right content
   1. Point you in the direction of the most popular tables or dashboards.
   2. Understand what dashboards exist for a certain business question or domain.
   3. Suggest tables to use to conduct analysis for specific questions.
3. Help generate queries
   1. Generate queries based on your metadata to calculate an insight.
   2. Recommend join paths between different datasets.

## Asking a Question

Ensure that the Ask AI Chatbot must be enabled for your organization. Once you have access, you can chat with AI Assistant directly in the app. Your personal history will be available in the left side bar.

<figure><img src="/files/QhqLky4YHaXPv6SsP5ui" alt="Select Star AI Assistant - starting a chat"><figcaption><p>Select Star AI Assistant - starting a chat</p></figcaption></figure>

You can rate the responses and tell us how we're doing!

<figure><img src="/files/CB5HVXaSk51qz45xAXGa" alt=""><figcaption></figcaption></figure>

Conversation history is stored and analyzed to continually improve the performance and accuracy of Ask AI.

## Data Access

Select Star's AI capabilities leverage OpenAI's and Anthropic's GPT technology. OpenAI and Anthropic are SOC-2 compliant vendors, and our security team reviews their SOC-2 report annually.

When using AI features of Select Star, the following data may be shared with OpenAI and Anthropic:

* **Metadata:** Names of data objects (such as schemas, tables, columns, dashboards), object descriptions, object types, and the folders they belong to.
* **Select Star Metadata Analysis:** Popular SQL queries, popularity scores, table creation statements, and names of data object sources/destinations (i.e., data lineage information).
* **User Metadata:**
  * **Select Star Users and Teams:** Email and first and last name of users in Select Star, as well as team names and descriptions.
  * **Data Source Users:** Names, emails, or identifiers of users associated with data sources - for example, table/dashboard owners or users running queries.
* **Select Star Users and Teams:** Email and first and last name of users in Select Star, as well as team names and descriptions.
* **Documentation:** User-written documentation (docs, metrics, glossary) within Select Star.
* **Chat Feedback:** Any user feedback (ratings or text) from using the AI features.
* *(Optional)* **Data values:** If your organization has authorized data access and enabled features like AI Data Querying, Select Star may execute SQL queries on behalf of users and process the resulting data.

To be explicit, the following data is not shared outside of Select Star:

* User or account information.
* Data in queries classified as sensitive, according to PII tags, will be scrubbed or masked.
* Data values — Select Star does not access, store, or transmit data values unless data access has been explicitly granted and features such as AI Data Querying have been enabled.

**OpenAI and Anthropic do not store or use any of the above data for training.**

Select Star’s AI features (such as Ask AI) can be disabled if customers prefer not to have any metadata sent to OpenAI or Anthropic.


# MCP Server

## Overview

The Select Star MCP Server enables MCP Clients like Cursor and Claude Desktop to access your Select Star MCP tools. Using the Model Context Protocol (MCP), your AI assistant can search for tables, dashboards, and columns, retrieve detailed asset information, and help you understand your data lineage.

## What is MCP?

Model Context Protocol (MCP) is a standard that allows AI assistants to connect to external systems and data sources securely. When you configure the Select Star MCP Server, your AI assistant gains the ability to:

* Answer questions about your data catalog
* Find relevant tables and columns for your analysis
* Provide context about data sources and their relationships
* Help you understand data lineage and dependencies
* Support dbt development workflows with schema discovery and production usage insights

## Getting started

#### Prerequisites

Before you can use the Select Star MCP Server, you'll need:

1. [Select Star API Token](/select-star-api/authentication)
2. MCP Client (Claude Desktop, Cursor IDE, etc.)
3. When using `mcp-remote@latest` in the configuration: Node.js (version 18 or higher)
   * Download and install Node.js from [nodejs.org](https://nodejs.org/en/download)
   * Follow the installation guide for your operating system
   * Verify installation by running `node --version` in your terminal

#### Quick Setup

The Select Star MCP Server works with any MCP Client. Here's how to configure it:

#### Cursor IDE

In Cursor, you need to configure the MCP server in your settings:

1. Open Cursor Settings
2. Navigate to "Tools & Integrations" → Click on "New MCP Server"
3. Add the following configuration to your MCP servers section:

```json
{
  "mcpServers": {
    "select-star": {
      "url": "https://mcp.production.selectstar.com/mcp",
      "headers": {
        "Authorization": "Bearer <your_actual_api_token_here>"
      }
    }
  }
}
```

**Important:** Replace `<your_actual_api_token_here>` with your Select Star API token.

#### Claude Code

For Claude Code, run this command in your terminal to add within the user scope:

```bash
claude mcp add --transport http select-star https://mcp.production.selectstar.com/mcp --header "Authorization: Bearer <your_actual_api_token_here>" -s user
```

**Important:** Replace `<your_actual_api_token_here>` with your Select Star API token.

#### Claude Desktop

For Claude Desktop, you need to modify the configuration file. You can follow the guide [here](https://modelcontextprotocol.io/quickstart/user) or:

1. Open Claude Settings
2. Navigate to "Developer" → Click on "Edit Config"
3. Add the following configuration to your MCP servers section in `claude_desktop_config.json`:

```json
{
  "mcpServers": {
    "select-star": {
      "command": "npx",
      "args": [
        "-y",
        "mcp-remote@latest",
        "https://mcp.production.selectstar.com/mcp",
        "--header",
        "Authorization:${SELECT_STAR_TOKEN}"
      ],
      "env": {
        "SELECT_STAR_TOKEN": "Bearer <your_actual_api_token_here>"
      }
    }
  }
}
```

**Important:** Replace `<your_actual_api_token_here>` with your Select Star API token.

#### Other MCP Clients

This server follows the standard MCP protocol and should work with any MCP Client. Adapt the configuration format to match your client's requirements while using the same connection details.

## Available Tools

We're continuously expanding the MCP Server's capabilities. If you need functionality that isn't currently available, please reach out to us at <support@getselectstar.com> - we'd love to understand your use case and may be able to prioritize it in our roadmap.

### `search_metadata`

Search across all data catalog metadata with advanced filtering capabilities.

### `get_asset_details`

Retrieve comprehensive information about a specific data asset.

### `get_table_info`

Retrieve detailed table schema and usage information.

### `traverse_lineage`

Explore data lineage relationships and dependencies.

### `get_asset_queries`

Retrieve data warehouse queries connected to tables or columns.

### `search_semantic_models`

Search for semantic models in the data catalog.

### `get_semantic_model_details`

Retrieve complete details about a specific semantic model, including YAML content.

## Troubleshooting

#### Authentication Failed

1. Verify your API token is correct
2. Ensure the token is properly formatted with "Bearer " prefix

#### Node.js Issues

1. Ensure Node.js version 18 or higher is installed
2. Try running `npx --version` to verify npx is available
3. Clear npx cache with `npx --cache-dir clear` if you encounter package issues

#### Connection Issues

1. Verify your internet connection
2. Check if your firewall allows connections to `https://mcp.production.selectstar.com/mcp`
3. Try running the npx command manually in your terminal to test connectivity


# Semantic Models

{% hint style="info" %}
Semantic Models is currently in **Beta**. We're actively improving the feature based on user feedback. If you encounter any issues or have suggestions, please contact us at <support@getselectstar.com>.
{% endhint %}

## Overview

Semantic models bridge the gap between technical metadata and business context. While traditional data catalogs show table names, column types, and lineage, semantic models define the **entities**, **metrics**, **relationships**, and **business logic** that represent how your organization actually uses data.

By creating semantic models in Select Star, you enable:

* **AI-powered insights**: Ask AI can reference business concepts like "customer," "ARR," or "churn rate" with proper context, generating accurate SQL queries
* **Consistent metrics**: Define metrics once and use them across your organization
* **Business context**: Add descriptions, synonyms, and sample business questions that help both humans and AI understand your data
* **Integration**: Export to tools like Snowflake Cortex Analyst or use with Select Star MCP for AI coding assistants

**Semantic models work with all data warehouses and BI tools supported by Select Star.**

## Quick Start

Want to create your first semantic model in 5 minutes? Here's the fastest path:

1. **Choose your source**:
   * **Have dashboards?** Go to your BI tool (Power BI, Tableau, etc.) and select 1-3 dashboards that represent a business domain
   * **No dashboards?** Go to your data warehouse and select the core tables for a business area (e.g., orders, customers, payments)
2. **Generate the model**:
   * Select your assets → Click **AI Assist** → **Generate Semantic Model**
   * Give it a name → Click **Generate**
   * Wait for email notification (usually 2-5 minutes)
3. **Enable in Ask AI**:
   * Open your semantic model → Go to **Settings** tab
   * Toggle **Enable Ask AI** to ON
   * Add custom instructions if needed (optional)
4. **Start asking questions**:
   * Go to Ask AI → Select your semantic model
   * Ask business questions like "What are our top customers by revenue?"

That's it! For detailed instructions and advanced features, continue reading below.

## Getting Started

### Prerequisites

Before creating semantic models:

1. **Select Star AI features** must be enabled for your organization
2. **Permissions**:
   * **Admins and Data Managers**: Can create, edit, and delete semantic models
   * **Viewers**: Can view semantic models but cannot make changes
3. **Source data**: Have at least one dashboard or table in Select Star to use as a source

### Creating a Semantic Model

You can create semantic models from dashboards or tables. Both follow the same basic process but differ in their starting point:

| **From Dashboards**                                   | **From Tables**                           |
| ----------------------------------------------------- | ----------------------------------------- |
| Best for capturing existing BI logic                  | Best for new semantic layers              |
| Uses **relevant lineage** (only contributing columns) | Includes all columns from selected tables |
| Ideal when metrics already exist in BI tools          | Ideal for full control over entities      |

#### Creation Steps

The creation process is the same regardless of source:

**For Dashboards:**

1. Navigate to your BI tool (Power BI, Tableau, etc.)
2. Select one or more dashboards
3. Click **AI Assist** → **Generate Semantic Model**

![Create from Dashboards](/files/w0zS7GtZBejCRYPCmmO7)

**For Tables:**

1. Navigate to your data warehouse (Snowflake, BigQuery, etc.)
2. Select one or more tables
3. Click **AI Assist** → **Generate Semantic Model**

![Create from Tables](/files/OD491LalQOnj09rjIyiB)

**Then for both:**

4. Enter a **Display Name** for your semantic model
5. Click **Generate**

![Name your Semantic Model](/files/B28Cx4lkidZA2mbyaRkG)

#### What Gets Generated

Select Star automatically creates:

* **Entities**: Logical tables based on your source data
* **Relationships**: Inferred from primary/foreign keys in your ERD
* **Metrics**: From dashboard calculations, existing metrics, and downstream dashboards
* **Example Queries**: Generated from table structure and common patterns

All components are AI-generated and can be edited after creation.

#### Generation Time

* **Standard models**: Typically complete within a few minutes
* **Large or complex models**: May take up to an hour if:
  * The semantic model includes many tables
  * Descriptions need to be generated for columns
  * Metrics need to be calculated
  * Example queries need to be created

You'll receive an email notification when generation is complete. You can monitor progress from the Semantic Models list page.

## Working with Semantic Models

### List View

Find all semantic models in your organization with their status, descriptions, and creation details. Search and filter to locate specific models quickly.

![Semantic Models List](/files/SfWvcldum0PRhVrtMKgi)

### Overview Tab

View the semantic model's purpose, business questions it answers, source assets, and ownership details. Track changes in the history panel.

![Overview Tab](/files/fvzpOTMjtqR3FA6Hd0E1)

### Entities Tab

Browse all logical tables (entities) in your model with their fields, data types, and business descriptions.

![Entities Tab](/files/KejC5qPNnOu3tS3ICAUA)

#### Adding Entities

To add new entities to your semantic model:

1. Navigate to the **Relationships** tab
2. Click the **Add Entity** button
3. Select a data source (e.g., Snowflake)
4. Choose a table from the dropdown
5. Select which columns to include:
   * Check **Select All** to include all columns
   * Or select individual columns
6. Click **Add Entity**

![Add Entity](/files/Lbxn8IKz4QeAa9Zjsk6h)

The new entity will be added with:

* All selected fields
* Inferred data types
* Descriptions from the source table (or AI-generated if none exist)

#### Editing Entities

You can modify entities through the **Relationships** tab:

* Click the menu (⋮) on an entity card
* Select **Edit Fields** to change which columns are included

#### Removing Entities

To remove an entity:

1. Navigate to the **Relationships** tab
2. Click the menu (⋮) on the entity card
3. Select **Delete Entity**
4. Confirm the deletion

{% hint style="warning" %}
**Important**: Deleting an entity will also delete all related metrics and example queries that reference that entity. This action cannot be undone.
{% endhint %}

![Delete Entity Warning](/files/wIHkBgfBLG1z1TXxUx2y)

### Relationships Tab

The Relationships tab shows how entities join using primary key and foreign key relationships.

![Relationships Tab](/files/0iyHYV62tXOQeUlbgRB0)

Visualize connections between tables and manage how entities relate to each other. Add new entities or edit existing relationships as needed.

### Metrics Tab

View business metrics with their definitions, synonyms, and SQL expressions. By default, semantic models include the top 10 most relevant metrics to ensure optimal AI performance.

![Metrics Tab](/files/DtVx0sgs0oAU2xvNgUTj)

### Example Queries Tab

See popular business questions with their SQL answers. These queries help Ask AI understand typical questions and demonstrate what's possible with your data.

![Example Queries Tab](/files/1O2Pm6BcASj7Bp46nudR)

### Settings Tab

Control how your semantic model integrates with Ask AI and external tools like Snowflake Cortex Analyst. Toggle Ask AI availability, add custom instructions for business rules, and configure Snowflake export options (YAML to stage or SQL view to database).

![Settings Tab](/files/dNzmBsBE157dM7aJf6XW)

## YAML Specification

### Viewing and Downloading YAML

Every semantic model has a YAML representation following the Snowflake Cortex Analyst specification.

To view the YAML:

1. Click the **YAML** button in the top-right corner of any semantic model page
2. The YAML modal displays the complete model definition

![YAML View](/files/WyvXoxXVYkEUePlc0bGX)

**YAML Contents:**

* Model name and description
* Business questions
* Source assets and tables
* Entity definitions (dimensions, time dimensions, facts)
* Relationships between entities
* Metrics with expressions
* Example queries with SQL

**Download Options:**

* Click the download icon to save the YAML file
* Filename format: `{model_name}_{date}.yaml`
* The downloaded file can be:
  * Used with Snowflake Cortex Analyst
  * Imported into other tools
  * Version controlled in Git
  * Shared with your team

### YAML Structure

The YAML follows this structure:

```yaml
name: Model Name
description: Model description and business context

entities:
  - name: EntityName
    description: Entity description
    base_table:
      database: DATABASE
      schema: SCHEMA
      table: TABLE_NAME
    dimensions:
      - name: column_name
        synonyms: [alternative_name, ...]
        description: Column description
        expr: column_name
        data_type: TYPE
        sample_values: [...]
    time_dimensions:
      - name: date_column
        ...
    facts:
      - name: numeric_column
        ...

relationships:
  - name: relationship_name
    left_table: EntityA
    right_table: EntityB
    relationship_columns:
      - left_column: column_a
        right_column: column_b

metrics:
  - name: metric_name
    synonyms: [...]
    expr: SUM(EntityName.column_name)

example_queries:
  - name: query_name
    question: Business question in natural language?
    sql: SELECT ... FROM ...
```

## Integrations

### Ask AI

Semantic models enhance Ask AI's ability to understand your data and generate accurate queries.

#### How Semantic Models Help Ask AI

When a semantic model is enabled in Ask AI:

1. **Business Context**: Ask AI understands business terms and metrics, not just table names
2. **Accurate Joins**: Relationships ensure correct table joins
3. **Consistent Metrics**: Metrics are calculated the same way every time
4. **Sample Questions**: Example queries guide what's possible

#### Using Semantic Models in Ask AI

When semantic models are enabled for Ask AI, they automatically enrich the AI's context. The AI agent intelligently selects the most appropriate semantic model(s) to answer your questions.

![Ask AI with Semantic Models](/files/cC9QCrnE9EX1RhLqmYaq)

**You can also manually select a semantic model** from the dropdown if you want to focus on a specific business domain.

Ask AI will:

* Automatically use relevant semantic models based on your question
* Apply business entities, relationships, and metrics
* Generate accurate SQL using your business terminology
* Suggest follow-up questions from example queries

**Benefits:**

* **Smarter AI**: Understands your business terms, not just table names
* **Accurate queries**: Correct joins and metric calculations
* **Faster answers**: Pre-defined context speeds up query generation

### Select Star MCP

Semantic models are also available through the [Select Star MCP Server](/features/mcp-server), enabling AI coding assistants like Cursor and Claude Desktop to:

* Discover and understand your semantic models
* Reference business entities and metrics in code
* Generate queries using established business logic
* Access example queries for common patterns

**Example Use Case:**

You're writing Python code and ask your AI assistant: *"Using the Sales Analytics semantic model, write a query to calculate total revenue by customer segment"*

Your AI assistant will:

* Discover the "Sales Analytics" semantic model via MCP
* Use the defined "Revenue" metric
* Apply correct joins via relationships
* Generate: `SELECT customer_segment, SUM(revenue) FROM sales_model GROUP BY customer_segment`

You can also let the AI discover semantic models automatically by asking: *"What semantic models are available?"* or *"Write a query for customer revenue analysis"* and the AI will find and use the appropriate model.

This gives your AI assistant deep understanding of your business data context, improving code generation quality.

## Best Practices

**Create semantic models for:**

* Core business domains (Sales, Marketing, Finance)
* Data used by multiple teams
* Frequently asked business questions

**Choose dashboards when:**

* Metrics are already defined in your BI tool
* You want only relevant tables/columns

**Choose tables when:**

* Building a new semantic layer from scratch
* You need full control over entities

**Maintain quality by:**

* Adding business context to AI-generated descriptions
* Validating relationships match your business logic
* Using custom instructions to document business rules, date defaults, and key filters

## Troubleshooting

### Model Generation Taking Too Long

If your semantic model is stuck "In Progress":

* **Check email**: You'll receive a notification when complete (may take up to an hour for large models)
* **Review size**: Very large dashboards with many tables may take longer
* **Wait for metrics**: Metric generation can be time-intensive
* **Contact support**: If generation exceeds 2 hours, contact <support@getselectstar.com>

### Relationships Not Showing Correctly

If entities aren't connected in the ERD:

1. Verify primary keys are defined on your tables in Select Star's ERD
2. Check that foreign key relationships exist in your ERD
3. Manually add relationships if needed using the Relationships tab

### Snowflake Publish Failing

If publishing to Snowflake fails:

**For Stage publish:**

* Verify the stage exists: `SHOW STAGES LIKE 'stage_name' IN SCHEMA database.schema;`
* Check Select Star has WRITE permission on the stage
* Ensure stage name is fully qualified: `database.schema.stage`

**For Semantic View publish:**

* Verify database and schema exist
* Check Select Star has CREATE VIEW permission
* Ensure the credential used by Select Star has sufficient privileges
* Review Snowflake error logs for specific permission issues

### Metrics Missing or Incorrect

If expected metrics aren't appearing:

* **Check metric limit**: Only top 10 most relevant metrics are included by default
* **Review source assets**: Metrics come from your dashboards, tables, and existing metric definitions
* **Verify entity columns**: Ensure the columns used in metric expressions exist in your entities
* **Check complexity**: Very complex calculations may not be included

### Ask AI Not Using Semantic Model

If Ask AI isn't using your semantic model:

1. Verify **Enable Ask AI** is toggled on in Settings
2. Ensure you've selected the semantic model in the Ask AI interface
3. Check that your question relates to entities in the model
4. Try rephrasing using terms from entity and metric names

***

## Additional Resources

* [Ask AI Documentation](/features/ask-ai-chatbot) - Learn more about using semantic models with Ask AI
* [Select Star MCP](/features/mcp-server) - Use semantic models with AI coding assistants
* [Metrics Documentation](/features/documents/metrics) - Understand metrics in Select Star

Have questions or feedback? Contact us at <support@getselectstar.com>


# AI Settings

Learn how to configure AI-powered features in Select Star. Enable AI-generated descriptions to automate documentation, customize chatbot data access.

Select Star has a number of AI powered features:

1. [AI generated documentation](/features/auto-documentation#ai-generated-documentation)
2. [Ask AI Chatbot](/features/ask-ai-chatbot)

Admins can configure these settings from **Admin > AI Settings** or directly under **Ask AI > AI Settings**.

Here, admins can

* Control the availability of all AI powered features via the **Enable AI Features** toggle
  * Separately enable/disable **AI Generated Descriptions** and the **Ask AI Chatbot**
* Configure the language of AI generated responses
* Enable AI to automatically document data assets when a new data source is added
  * This setting only applies when a new data source is added, not to existing data sources.
  * When enabled, descriptions will be generated for the top 20% of assets, by popularity and lineage

<figure><img src="/files/lYsuMiIcmLRFCKN4pET1" alt=""><figcaption><p>AI Settings</p></figcaption></figure>

## Ask AI Settings

{% hint style="info" %}
Note: The Ask AI settings apply to both the in-app and [Slack](/integrations/slack#answer-data-questions-beta) versions of Ask AI.
{% endhint %}

### Chatbot Prompt Customization

Admins can further provide custom instructions, enhancing the relevance and quality of the chatbot for your use cases.

Some examples include:

* Do not generate SQL queries unless explicitly asked
* Only return assets with the "Certified" tag
* If there are multiple matching resources, always return the most popular ones

Admins can add custom instructions and click **Save** to update them for all users. This applies immediately so you can quickly test out the impact with the Ask AI Chatbot.

<figure><img src="/files/y6rtrWOQsxb9Y5i4YeZS" alt=""><figcaption><p>Customize the Ask AI Chatbot prompt</p></figcaption></figure>

### Ask AI Chatbot Data Access

You can also restrict the access the Ask AI chatbot should have access to, in case there are data assets that are not relevant to most users.

By default the Ask AI chatbot has access to all data assets in your instance, including:

* Tables and columns
* Dashboards and charts
* Documentation: Pages, Metrics and Glossary Terms

<figure><img src="/files/6gla44ChX1coSCV5EONR" alt=""><figcaption><p>Default Ask AI Chatbot Data Access</p></figcaption></figure>

You can narrow down the available assets

* By data source (i.e. Tableau, or Snowflake)
* For data wareouse tools: by database and schema
* For BI tools: by folder
* For documentation in Select Star: **All Docs**, or by individual document asset

{% hint style="info" %}
Data Asset rules are treated as additive, so the chatbot will have access to all asset in any of the specified data sources, databases, schemas or folders.
{% endhint %}

{% hint style="warning" %}
If you remove **All** from the data access, be sure to add **All Docs** if you want the chatbot to have access to the pages, metrics and glossary terms that you have added in your instance.
{% endhint %}

You can further limit the Ask AI chatbot to use only assets tagged with specific tags i.e. **Certified**, or **Gold.** This applies on top of the available data assets.

In the below example, the chatbot will have access to

* any dashboard and chart in Tableau that have been tagged with **Gold** or **Certified**
* all pages, metrics, glossary terms in Select Star that have been tagged with **Gold** or **Certified**
* the tables and columns in the OLIST.DATASETS schema that have been tagged with **Gold** or **Certified**

<figure><img src="/files/qxctqgRu7CfyByWkumuY" alt=""><figcaption><p>Customized Data Access for the Ask AI Chatbot</p></figcaption></figure>

#### Things to Note

* Embeddings backing AI features are generated once a day. If you add a tag to a data asset, it will not be reflected in chatbot responses until the embedding is re-generated.
* If you are customizing data access with a tag filter, remember it applies to all assets that the chatbot has access to, including documentation, dashboards and tables.


# AI Agents (Private Beta)

{% hint style="warning" %}
AI Agents is currently available in Private Beta.

If you’d like to enable it for your organization, or if you have any questions or feedback, please reach out to Select Star Support or email <support@getselectstar.com>.
{% endhint %}

### Overview

AI Agents help you keep metadata consistent and organized by automatically applying actions to tables, dashboards, and columns. Using natural language, administrators can define automation rules that run daily to ensure metadata stays accurate and up to date.

<figure><img src="/files/zQGhiqfo4LiuDaHBwPd7" alt=""><figcaption></figcaption></figure>

Key Features

* Create tagging rules using natural language
* Intelligent filter validation with preview
* Tags, Owners, Collections management
* Bulk operations on tables, dashboards, and columns
* Advanced hierarchical filtering (e.g., "tag tables that have a column that has a Tag Y")
* Daily automatic execution

### Getting Started

#### Prerequisites

* **Private Beta access**: Contact Select Star Support
* **Admin role**: Required to create, view, enable/disable agents
* **AI features enabled**: In Admin Panel

#### Basic Workflow

1. Enable AI features in the Admin Panel.
2. Go to the AI Agents section in the Ask AI page.
3. Create a new agent by writing tagging instructions in natural language.
4. Review the preview to confirm the agent’s interpretation.
5. Type **"save"**, **"confirm"**, or a similar instruction to save the agent and schedule daily execution.
6. Once created, the agent is enabled by default and added to a queue for execution. The first execution may take some time if other agents are ahead in the queue. After the initial run, the agent will execute daily.
7. On the AI Agents page, you can enable, disable, or delete agents.

#### Available Actions

| Action                      | Description                      |
| --------------------------- | -------------------------------- |
| **Add tags**                | Apply tags to matching assets    |
| **Remove tags**             | Remove specific tags             |
| **Assign owners**           | Set business or technical owners |
| **Add to collections**      | Add assets to collections        |
| **Remove from collections** | Remove from collections          |

#### Filter Examples

You can define rules based on different criteria, such as names, dates, lineage, popularity, or existing tags:

<table><thead><tr><th width="196.03125">Filter Type</th><th>Example Command</th></tr></thead><tbody><tr><td><strong>By name</strong></td><td>Find columns named "customer_email"</td></tr><tr><td><strong>By date</strong></td><td>List tables created in the last 30 days</td></tr><tr><td><strong>By lineage</strong></td><td>Find tables with downstream lineage</td></tr><tr><td><strong>By popularity</strong></td><td>List dashboards with popularity > 80</td></tr><tr><td><strong>By data source</strong></td><td>Tag all snowflake tables as "Tested</td></tr><tr><td><strong>Hierarchical</strong></td><td>Find tables that have a column with "email" in name</td></tr></tbody></table>

#### Notifications

Asset owners are receive notification when AI Agents update assets they own. We also maintain the history of updates on the object overview page.

<figure><img src="/files/imgTBaWbWvfm33rbE6rG" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/40wkJv4uW94uphLFIWWa" alt=""><figcaption></figcaption></figure>

#### Limitations

* **Supported asset types**: tables, columns, and dashboards
* **Supported filter attributes**: name, data source type, popularity, created on, updated on, lineage (existence only), and hierarchical filters
* **Date filters**: `updated on` and `created on` are available only for tables and columns. Dashboards do not support these filters.
* **Lineage filters**: Can only check if upstream or downstream lineage exists or doesn't exist. Specific lineage questions (e.g., "tables connected to specific dashboard X") are not supported.
* **Filter validation**: The system validates filters and provides helpful error messages when unsupported filters are used. If a filter isn't recognized, the AI will suggest corrections or alternatives.

### Step-by-Step Tutorial (Tag Agent)

#### 1. Navigate to AI Agents

* Open the **Ask AI** page.
* Click on the **AI Agents** item.

<figure><img src="/files/HzqcsNCUAGHmYdosXWrM" alt=""><figcaption></figcaption></figure>

#### 2. Create a New Agent

* Click **Create Agent**.
* Enter a natural language instruction describing your tagging rule.

> List all tables with the name “Log”

<figure><img src="/files/cauooVFyuJ1CCtr9s087" alt=""><figcaption></figcaption></figure>

#### 3. If needed, update the filter.

> Update the filter to also include tables the name includes “analytics”.

**Note**: using “update” is important, otherwise the prompt will create a new filter with only the new instruction.

<figure><img src="/files/HXMGqV7eVuQSssiJmaLk" alt=""><figcaption></figcaption></figure>

#### 4. Add the tag

> Add "analytics" tag

<figure><img src="/files/0CCSR4IZFHjcNfMz7bQi" alt=""><figcaption></figcaption></figure>

#### 5. Save the agent

Type "save", "confirm", or a similar instruction to save the agent.

> Confirm

<figure><img src="/files/I01SnYTyAsGeol9XSItb" alt=""><figcaption></figcaption></figure>

Once the agent is running, you can track its progress in the list view and visit the chat to update the prompt.

<figure><img src="/files/KWJSwKcJGlp7H2BGmpsl" alt=""><figcaption></figcaption></figure>


# AI Admin Rating (Private Beta)

{% hint style="warning" %}
**AI Admin Rating is currently available in Private Beta.**

If you’d like to enable it for your organization, or if you have any questions or feedback, please reach out to Select Star Support or email <support@getselectstar.com>.
{% endhint %}

### Overview

AI Admin Rating provides a human-in-the-loop review workflow for Ask AI sessions. Admins can browse recent questions and answers, review statuses, ratings, and user feedback, add comments to coordinate with users or data owners, and monitor overall quality. Use it to evaluate **accuracy, clarity, and usefulness**.

#### Why this matters

* Increase trust in Ask AI with direct visibility into conversation quality and usage.
* Demonstrate value to leadership with real conversations and summary metrics.
* Make usage and quality visible so you can spot where users struggle and take action - then triage quickly and route fixes:
  * **Catalog/metadata improvements:** When metadata quality is low, data stewards should improve asset descriptions, tags, collections, and other metadata to ensure assets are well-documented and discoverable.
  * **Prompting guidance**: If a question is underspecified, suggest adding context or rephrasing.
  * **Ask AI behavior**: When issues are caused by incorrect prompt handling in Ask AI, flag them as *Select Star Issues* and route to Support.

#### Key features

* **Purpose-built evaluation**: Human-in-the-loop review of Ask AI conversations.
* **Status management**: Admins can set workflow statuses.
* **Evidence view**: Analyze the question, reasoning steps, generated SQL, and responses.
* **Quality tracking**: Overall metrics (total sessions, pass rate, average rating) to monitor progress. Use filters to identify patterns across sessions.

### Getting Started

1. Ensure your organization role is **Admin**.
2. Make sure AI features are enabled in your org.
3. Open `Ask AI` → `AI Admin Rating`.

#### Basic Workflow

Review the summary – check total sessions, Pass/Need Review counts, and average rating. Click a chat to see more details about a specific conversation:

<figure><img src="/files/wbosx52GIjQLvw0rli8G" alt=""><figcaption></figcaption></figure>

Analyze the conversation and leave a comment if you find any issues:

<figure><img src="/files/OwHcJiW32OvM6hIhNiam" alt=""><figcaption></figcaption></figure>

Take action – set a workflow status and add a comment if needed:

<figure><img src="/files/iD7utCDCLbW4KyzXC86c" alt=""><figcaption></figcaption></figure>

#### Ratings & Statuses

* `Pass` - Set automatically for ratings 4–5. Use when the answer is acceptable for business use.
* `Need Review` - Set automatically for ratings 1–3. Use when follow‑up is required to correct or clarify.
* `In Review` - Admin has picked it up to investigate. Signals ownership to avoid duplicate work.
* `Issue Found` - Admin flags an incorrect/low‑quality/risky answer. Add a note and coordinate with the user or data owner to resolve.
* `Select Star Issue` - Technical/model problem. Route to owner or Select Star Support.
* `Resolved` - Admin confirms the conversation is acceptable after fixes/re‑test.

#### FAQ

* **Can admins edit others' sessions?** Admins can view all sessions but cannot change ratings.
* **Can admins add comments to any sessions?** Yes. Admins can add comments on sessions they do not own.
* **Who gets notified when someone is mentioned?** Only admins receive notifications when they are mentioned. Mentions of other users do not trigger notifications.


# Data Discovery

Simplify data exploration with Select Star's data discovery platform. Explore these articles for more information.

This section contains the following articles:

{% content-ref url="/pages/-MgSX6Emz9KrSGbRN0xc" %}
[Where's my data?](/data-discovery/wheres-my-data)
{% endcontent-ref %}

{% content-ref url="/pages/-MgSXAiUZoCTigQ7NYhh" %}
[Where's my dashboard?](/data-discovery/wheres-my-dashboard)
{% endcontent-ref %}

{% content-ref url="/pages/-Mgw7LCSRDw7dq5FpYNa" %}
[How can I get the full context of this data?](/data-discovery/how-can-i-use-this-data)
{% endcontent-ref %}

{% content-ref url="/pages/-MgCExCX8lU1GX16wXiI" %}
[My dashboard looks off](/data-discovery/troubleshooting-broken-dashboards)
{% endcontent-ref %}

{% content-ref url="/pages/-MgSXHyUQ99m5C4Q89xI" %}
[I'm new to the team](/data-discovery/im-new-to-the-team)
{% endcontent-ref %}

{% content-ref url="/pages/-MgSXQXFWr7hbzXxkxrF" %}
[I have a data question](/data-discovery/i-have-a-data-question)
{% endcontent-ref %}


# Where's my data?

Effortlessly locate your data with Select Star's search and filters. Sort through datasets based on BI tools, customize views, and apply tags for efficient data exploration.

If you have any questions about where you data lives, you can always start at the [search bar](/features/search) at the top of the page.

Using search can quickly take you to a dataset even if you already know where it lives in your hierarchy. The search has filters which can narrow down results to a specific BI tool or type of data asset, but if you already know which tool the information you're looking for is in, you can sort through that information in different ways.

If you know you're looking for a data table or column, click the database it belongs to in the left sidebar.

![](/files/-MhtNHWgqnhsg3Aiv64F)

## Database Pages

From here, use `Cmd/Ctrl+F` or click the :mag: icon in the header to open a filter in the page. Type in a filter to narrow the results, or use the filters on the right side of the page.

![](/files/-Mibrx95lP2ZChiuPV2w)

Customize the information you see on this page by clicking the filter icon ![](/files/2qSgPtoqlNo1YzsBK7Tw)at the right side of the header.

![](/files/-MibsH-5iVQDZj1dhuvB)

### Filters

Change the **Data Type** in the dropdown to see all columns in a database. It will show all tables and views by default.

![](/files/-Mibsb83m4hrRPkNhPM4)

Check any of the [Category or Status Tag](/features/tags) boxes to show data with those tags applied.

![](/files/-MibsohlVszAKA0pitOv)

Find datasets which have no [Owners assigned](/data-management/data-ownership), or ones which belong to specific users or teams.

![](/files/-MibtLFz9KMVAvtKrF6J)

![](/files/-MhpElB2e84CughnNVEO)

See your most popular and least popular datasets, as well as those with no popularity, documentation, or tags. These filters can help you identify where you need to [add documentation](/data-management/add-documentation) or distinguish between excellent and subpar datasets [using tags](/data-management/tag-management).

![](/files/-MibtF63OBW_C8eF6l55)


# Where's my dashboard?

Find data in specific BI tools, apply category and status tags, and identify popular and well-documented dashboards using Select Star.

If you have any questions about where you data lives, you can always start at the [search bar](/features/search) at the top of the page.

Using search can quickly take you to a dataset even if you already know where it lives in your hierarchy. The search has filters which can narrow down results to a specific BI tool or type of data asset, but if you already know which tool the information you're looking for is in, you can sort through that information in different ways.

If you know you're looking for is in a BI Tool, check the left sidebar for...

[Looker](/learning-data/getting-started-looker) :point\_right: ![](/files/-MiRz6YZW_ekpc6X97zH)

[Tableau](/learning-data/getting-started-tableau) :point\_right: ![](/files/-MiRzQr73HNo7NHJ5ndO)

[Mode](/learning-data/getting-started-mode) :point\_right: ![](/files/-MiRzTxbj6xo2G0eRFom)

## Dashboard Pages

From here, use `Cmd/Ctrl+F` or click the :mag: icon in the header to open a filter in the page. Type something in to narrow the results.

![](/files/-MibpLwmsyeFRxfB9hQC)

### Filters

You can also use the **Filters** on the right side of the page.

![](/files/aOGvGxFM3xtBMndi2oe5)

For Mode Reports and Looker Dashboards, you can filter by reports created by a specific user.

![Click +Add Filter to choose a user.](/files/-Mibq4jLQ0lWedVDGpY6)

In Looker Explores and Tableau Dashboards, you can change the **Data Type** filters.

![](/files/-MiS0aYR08YM29o9ODs8)

![](/files/-MhtMAv-hHgbXEnqD64Y)

Check any of the [Collections](/data-management/collections) or [Tags](/features/tags) filters to show data with those tags applied.

<figure><img src="/files/g5jTKGv9nK1Tvjeu8irv" alt=""><figcaption></figcaption></figure>

See your most popular and least popular dashboards, as well as those with no popularity, documentation, or tags. These filters can help you identify where you need to [add documentation](/data-management/add-documentation) or distinguish between excellent and subpar dashboards [using tags](/data-management/tag-management).

![](/files/-MibR_eUVoSH7Loh6aPq)




---

[Next Page](/llms-full.txt/1)

