Using the Enhanced Web Crawler for Your AI Agent's Knowledge Base
What Is the Enhanced Web Crawler?
The Enhanced Web Crawler is a powerful tool within your CRM's AI Agent training system. It is designed to gather comprehensive information from websites by simulating real user interactions. Unlike basic crawlers that only read static content, this advanced engine opens tabs, expands accordions, scrolls through pages, and triggers lazy-loaded elements to capture hidden or dynamically loaded text. This ensures your AI bot has access to a much richer dataset for answering questions accurately.
Key Benefits
Deeper Content Extraction
The crawler captures up to 50% more on-page content from modern websites, including those built with frameworks like React, Vue, or Angular. It reads text within accordions, modals, tabs, and infinite-scroll sections.
Intelligent Dynamic Content Handling
It uses over 12 parallel detection strategies to quickly identify and extract meaningful content such as product descriptions, testimonials, pricing details, and contact information. The crawler avoids disruptive actions like form submissions or filter changes to ensure safe operation.
Advanced Link Discovery
The tool supports recursive sitemap crawling, processes compressed sitemap files (like .xml.gz), and uses navigation guards to stay within intended website boundaries. It discovers links hidden behind dynamic elements and deduplicates URLs while preserving descriptive link text.
Universal Website Support
It works with any website type, from static HTML to complex single-page applications. Detailed metrics provide full observability into crawl performance, including processing time, interactions, content length, and memory usage.
How to Use the Enhanced Web Crawler
Step 1: Access Knowledge Base
Navigate to AI Agents from your sub-account dashboard. Click on the Knowledge Base tab. Create a new Knowledge Base or edit an existing one. Click the + Add Source button and select Web Crawler.
Step 2: Enter Domain Information
Choose your domain type to define the crawl scope:
- Exact URL: Crawls only a specific webpage.
- All URLs with the Path: Crawls all pages within a specified path.
- All URLs in this Domain: Crawls every page on the domain.
Enter the URL and click Extract Data.
Step 3: Select URLs for Training
After crawling completes, click View All Pages. Select all URLs or individual ones using the checkboxes. Click Train Bot to add the selected content to your AI's knowledge base.
Frequently Asked Questions
What content does it capture?
It captures significantly more website content than before, including testimonials, features, contact details, and service descriptions that were often missed.
Is it reliable?
Success rates have improved dramatically across various site types, making ingestions far more reliable.
Does it require configuration?
No. Multiple parallel detection strategies automatically find key sections like hero content, pricing tables, and team bios.
Can it read interactive content?
Yes. It expands accordions, navigates tabs, and reveals lazy-loaded sections to capture full content.
What about structured data?
It extracts much more structured data (e.g., business hours, services), giving your AI a richer understanding of your business.
Will it submit forms or click checkout buttons?
No. The safe-interaction engine ignores form elements to prevent accidental submissions.
What about login-gated content?
It only crawls publicly accessible content. Private or login-protected data is not included.