> ## Documentation Index
> Fetch the complete documentation index at: https://docs.usefini.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Sources

> Connect links, files, Google Drive, Notion, Confluence, and Zendesk help centers as raw inputs that flow through Review and into Articles.

Sources is one of the main ways knowledge enters Fini's [Knowledge](/en/knowledge/overview) model. Anything you ingest here (help center URLs, uploaded PDFs, Notion pages, Zendesk articles, Confluence spaces) becomes a candidate for knowledge. It doesn't answer customers directly. It feeds the same review / publish workflow that lands in [Articles](/en/knowledge/articles), which is what the agent actually retrieves from at runtime.

<Frame>
  <img src="https://mintcdn.com/fini/zef_RDWlJqADKDlS/images/en/knowledge/sources/list.png?fit=max&auto=format&n=zef_RDWlJqADKDlS&q=85&s=82e165427ea231b4a8b56b1f6e3a1cf2" alt="Sources page in the Fini Demo workspace showing a Links source summary card, search and Read status filters, the Add Sources button, and a table of Venmo help-center links already added to knowledge" width="1680" height="1050" data-path="images/en/knowledge/sources/list.png" />
</Frame>

<Info>
  **Why Sources isn't the source of truth.** A help center is written for humans browsing a website, it has marketing language, outdated pages, contradictions between articles, and gaps where the answer is "ask support." Pointing an agent directly at it inherits all of that. Sources contributes raw material; Articles is the curated, deduplicated, conflict-resolved version your agent actually uses. See [Knowledge](/en/knowledge/overview) for the full model.
</Info>

The quality of what eventually reaches Articles is bounded by what's in Sources. A thin or stale set of sources produces thin or stale articles. A well-maintained Sources catalog is the biggest top-of-funnel lever you have over agent quality.

## Source types

Fini supports six source types out of the box. The difference between them is how the content gets in: some you upload directly, some you pull from a URL, some come over a connected OAuth integration.

| Type | What it accepts | How content arrives |
| - | - | - |
| **Links** | Public URLs (single, crawled from a root, or a sitemap) | Fini crawls and re-crawls on demand. |
| **Files** | PDF, DOC, DOCX, XLS, XLSX, CSV, HTML, Markdown, JSON, YAML | You upload directly; Fini parses + indexes. |
| **Google Drive** | Docs, Sheets, PDFs from a connected Google account | OAuth picker; you select files in the modal. |
| **Notion** | Pages and databases from a connected Notion workspace | OAuth picker. |
| **Confluence** | Spaces and pages from a connected Confluence site | OAuth picker. |
| **Zendesk help center** | Help center articles from a connected Zendesk subdomain | OAuth + subdomain. |

<Warning>
  **Zendesk shows up twice in the navigation** for a reason. Under **Sources**, Zendesk means "ingest articles from your Zendesk help center as agent knowledge" (read-only). Under **[Deploy → Zendesk](/en/deploy/zendesk)**, Zendesk means "post agent replies on incoming Zendesk tickets" (read + write on conversations). They use different OAuth scopes and are configured independently. Connecting one does not connect the other.
</Warning>

## Adding sources

Click **+ Add Sources** in the top-right of the page. The **Add Source** dialog opens with a tile per source type ("Popular" badges flag the most-used ones); pick the type you want, and Fini hands you off to the type-specific add flow.

<Frame>
  <img src="https://mintcdn.com/fini/zef_RDWlJqADKDlS/images/en/knowledge/sources/add-sources-modal.png?fit=max&auto=format&n=zef_RDWlJqADKDlS&q=85&s=0a7975b8c90bfc29624e9400de43b4f8" alt="Add Source dialog with tiles for Links, Files, Google Drive, Notion, Confluence, Zendesk" width="1680" height="1050" data-path="images/en/knowledge/sources/add-sources-modal.png" />
</Frame>

<Tabs>
  <Tab title="Links">
    The most common source type. The **Add link sources** dialog has three modes:

    * **Use individual links**: paste one or more URLs (newline-separated). Each becomes its own document. Use when you want exact control over what's added.
    * **Crawl parent link**: paste a single root URL. Fini fetches the page, finds every linked page on the same domain, and lists them for you to review and trim before submitting. Use when you want everything under a marketing or help-center domain in one shot.
    * **XML Sitemap**: paste a `sitemap.xml` URL. Fini parses the sitemap and lists every URL inside. Use when the site exposes a sitemap (which is how most modern docs and blog sites publish their canonical URL list).

    There's a per-batch cap on how many links you can add at once (configurable via your workspace; default is 10). If a crawl returns more than the cap, the modal asks you to trim the list before submitting.

    <Tip>
      Start with **XML Sitemap** if the source publishes one. It's faster, more accurate, and gives you the publisher's canonical URL list without missing pages or pulling in noise pages (privacy policy, etc.) that **Crawl parent link** sometimes includes.
    </Tip>

    <Frame>
      <img src="https://mintcdn.com/fini/zef_RDWlJqADKDlS/images/en/knowledge/sources/add-link-source.png?fit=max&auto=format&n=zef_RDWlJqADKDlS&q=85&s=e3391f10256a828000e68d8f25679fd3" alt="Add link sources dialog with three mode buttons: Use individual links, Crawl parent link, XML Sitemap" width="1680" height="1050" data-path="images/en/knowledge/sources/add-link-source.png" />
    </Frame>
  </Tab>

  <Tab title="Files">
    Drag-and-drop or browse to upload one or more files. Accepted formats: `.pdf`, `.doc`, `.docx`, `.xls`, `.xlsx`, `.csv`, `.html` / `.htm`, `.md` / `.markdown`, `.json`, `.yaml` / `.yml`.

    Each file becomes one or more documents (long files are chunked). The original file is preserved, so you can swap it later by uploading a new version with the same name.
  </Tab>

  <Tab title="Connected stores">
    For Google Drive, Notion, Confluence, and Zendesk, the flow is the same:

    <Steps>
      <Step title="Pick the source type">
        Click the source type tile in the Add Sources modal.
      </Step>

      <Step title="Authorize">
        If your workspace hasn't authorized that source yet, the OAuth flow runs first. For Zendesk, you also enter your subdomain.
      </Step>

      <Step title="Pick content">
        Once authorized, Fini lists the available content (Drive files, Notion pages, Confluence spaces, Zendesk articles).
      </Step>

      <Step title="Submit">
        Pick the items to ingest and click submit. Fini fetches and indexes them.
      </Step>
    </Steps>

    <Note>
      You only authorize the OAuth connection once per source. Subsequent visits to the modal go straight to the picker.
    </Note>
  </Tab>
</Tabs>

## The Sources page

The page shows every document Fini has indexed across all your sources. Summary cards at the top group the workspace by source type and health, for example **Links** with a count and **Healthy** status. Click a card to focus the table on that source type.

Under the cards, the toolbar gives you:

* **Search sources...** — search by source title.
* **Read status** — filter by whether a source has already been added into knowledge.
* **+ Filter** — add extra filters when you need to narrow the table further.

Each row shows the source title, source domain or filename, knowledge/read status, and age. The row actions let you inspect the parsed content, open the original source, refresh it, or delete it.

<Tip>
  **Read status = promotion state.** If you've just ingested a help center and want to see what has not yet made it into knowledge, use the **Read status** filter. That's your candidate list for article generation.
</Tip>

## Refreshing sources

Sources are stable until you change them. They don't auto-refresh on a schedule by default; if your help-center articles change, Fini keeps working from the version you last ingested unless you explicitly refresh.

To refresh:

* **One source**: select the row, open its actions menu, pick **Refresh**.
* **Several sources**: select multiple rows from the table and bulk-refresh.
* **All sources of a type**: select the type tab, select-all, refresh.

When a refresh produces a different version of a document, the **Diff modal** shows what changed (added/removed/edited lines) so you can confirm the new content reads correctly before it's indexed. Once indexed, the changes flow downstream into Review (as proposed updates to any Articles that reference this source).

When generated or updated Articles reuse a passage from a source, Fini preserves inline links and image references that belong to that passage. Links and images from unused passages are not carried over.

<Tip>
  Re-crawl your public marketing or help-center pages whenever you ship a substantial copy update. Fini has no way to know the underlying page changed otherwise — and any Article derived from a stale Source will stay stale until you refresh.
</Tip>

## Training preferences

Two language settings control how Fini processes your sources:

* **Language**: pick the language of your knowledge base. Defaults to **English**. If your sources are in multiple languages, switch to **Auto-detect** (Fini detects each document's language) or pick a specific non-English language.
* **English-only mode**: a faster path for English-only knowledge bases that skips per-document language detection.

These preferences are available every time you do a new source ingestion.

## How sources reach an agent

Sources live at the **workspace level**: anything you upload or connect is in the workspace's raw input pool. From there, the path to a customer-facing answer goes through two more layers:

<Steps>
  <Step title="Ingest into Sources">
    Connect the URL, upload the file, or pick from the OAuth integration. The content is indexed at the workspace level.
  </Step>

  <Step title="Promote to Articles">
    Either author an Article manually that references the source, or use [Magic Articles](/en/knowledge/magic-articles) to generate articles from your Sources in bulk. Either way, the content enters the same review / publish workflow before it becomes canonical knowledge.
  </Step>

  <Step title="Approve in Review">
    A teammate vets the draft in [Review Queue](/en/knowledge/review) — checking for conflicts with existing articles, gaps, and quality issues. On approval, the article goes live.
  </Step>

  <Step title="Attach to an agent">
    In [Articles](/en/knowledge/articles), use the top-right selector to attach the article (or its parent folder) to the agents that should retrieve from it. Different agents can have different knowledge bases drawn from the same Sources pool.
  </Step>
</Steps>

<Note>
  Adding or removing a source from your workspace doesn't automatically push changes to any agent. The downstream Articles are what agents retrieve from. Refresh the source, let the proposed updates flow into Review, approve them, and the changes propagate.
</Note>

## Why an answer is wrong or stale

Sources are upstream of every agent answer, so when something's off, the cause is usually one of these. In rough order of likelihood:

<AccordionGroup>
  <Accordion title="The source has been updated since you last ingested it" icon="rotate">
    Refresh the source. If you've never refreshed since the original ingest and the underlying page has changed, the Article derived from it is still working off the old version.
  </Accordion>

  <Accordion title="The source never made it into an Article" icon="circle-xmark">
    Sources alone don't answer customers — Articles do. Open [Magic Articles](/en/knowledge/magic-articles) and generate from this source, or author an Article manually. Check Review for any drafts sitting unapproved.
  </Accordion>

  <Accordion title="The relevant Article isn't attached to the responding agent" icon="robot">
    Open [Articles](/en/knowledge/articles), switch the top-right selector to the responding agent, and confirm the parent folder is attached.
  </Accordion>

  <Accordion title="The source format stripped the answer at parse time" icon="file-pdf">
    Click **Preview content** on the source's row to see exactly what Fini extracted. Some PDFs (especially scanned image-only PDFs) lose text during ingest — if the relevant passage isn't in the preview, it never reached Articles. Fix by re-exporting the file as a text-extractable PDF or uploading it as Markdown.
  </Accordion>

  <Accordion title="The crawl missed the page that has the answer" icon="link-slash">
    Re-run the crawl in **Sitemap** mode if the site exposes one, or paste the specific URL in **Individual** mode.
  </Accordion>

  <Accordion title="A conflict between sources wasn't resolved in Review" icon="triangle-exclamation">
    Two sources disagreed, and Review surfaced the conflict but it was approved without resolution. Open the article in [Articles](/en/knowledge/articles), fix the canonical answer, and resubmit through Review.
  </Accordion>

  <Accordion title="The agent's language doesn't match the source's language" icon="language">
    Check the language preference on both the source and the agent. An agent set to English-only won't reliably retrieve from non-English-derived Articles.
  </Accordion>
</AccordionGroup>

## Deletion

Deleting a source removes it from the index immediately. Any Articles derived from that source are **not** automatically deleted, they continue to serve customers from their already-captured content. If you want the downstream Article gone too, delete it explicitly in [Articles](/en/knowledge/articles).

The original file (for **Files**) remains in your workspace's storage; the link metadata is purged. Bulk-deletion via the table multi-select is supported.

<Warning>
  Deletion is **irreversible** for ingested data. To restore, re-ingest from the original source.
</Warning>


## Related topics

- [List sources](/en/api-reference/list-sources.md)
- [Get source](/en/api-reference/get-source.md)
- [Refresh sources](/en/api-reference/refresh-sources.md)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.