Skip to main content

Knowledge Foundation

Every surface and every tier depends on the same thing underneath: real, current, trusted content. Get ingestion right and everything downstream gets easier.


Getting Started Fast: Auto-Create Agent

The fastest path from zero to a working demo or Crawl-tier deployment: provide a website URL and a few setup choices, and ai12z automatically ingests the content, builds the knowledge base, applies branding, and creates the initial experience: a configured agent ready to review, test, and publish in minutes. See Creating an Agent for the full flow.


Keeping Content Current

Two ingestion paths, and most clients end up using both:

  • Event-driven CMS connectors: for clients on a supported platform, ai12z listens for publish events, so content added, updated, or deleted stays in sync automatically without a manual re-ingest. Real connectors exist today for Drupal, WordPress, Contentstack, Sitecore, Magnolia, Optimizely, Umbraco, Sitefinity, Agility, Kentico, Kontent.ai, Storyblok, and Salesforce Knowledge, plus AWS S3 and SharePoint for document/file sources. See Introduction to Connectors for the full list and setup.
  • Scheduled web scraping: for a client without a supported connector, ai12z falls back to crawling the public site on a schedule (daily, weekly, or monthly) or on demand. See Website Ingestion.

Always check first whether the client's platform has a supported connector: it's more reliable than scraping and removes an entire category of problems. When scraping is the only option, check the client's site for Cloudflare or similar bot protection before you scope the timeline. Some sites block crawlers outright, and getting a client's IT or CDN team to whitelist the crawler can take longer than the ingestion itself. See Whitelisting the ai12z Crawler for the exact IP, user-agent, and a ready-to-forward request template.

Teams can also ingest specific URLs, product pages, support articles, documentation, PDFs, catalogs, spec sheets, and compliance documents directly, see Content Ingestion Overview.

What to tell the client: whichever path fits their stack, the result is the same promise: the AI's answers stay grounded in what's actually published, not a stale snapshot from setup day.


Multilingual Content

Detecting and replying in a visitor's language is automatic (see Multimodal and Multilingual by Default), but citing the right page in that language is not: it requires that page to actually be ingested. For a client with real per-language content (e.g. a Canadian site with separate French and English pages), ingest each language separately rather than in one combined run, so a French question cites the real French page instead of the English one. Website Ingestion supports filtering an ingestion run by detected language for exactly this reason.

If the client only has content in one language but needs to answer in several, the AI can still translate its answer on the fly, but there's no localized page to cite, so set that expectation during scoping rather than after launch.


Image AI

Ingestion isn't just text. Image AI analyzes images during ingestion and generates descriptions that support search, accessibility, and AI-generated answers, and when an answer has a relevant supporting image, ai12z can return both the text and the image together, matched automatically rather than requiring anyone to manually tag which image goes with which content.