The company's first support AI had a trust problem. It answered instantly and confidently — and often from thin air. Wrong refund windows, outdated plan limits, policies that had changed months ago. Every wrong answer created a follow-up ticket and a customer slightly less willing to believe the next answer. The team was weeks from switching it off.
40-person B2B SaaS · Support AI across web widget and internal Slack · Forge RAG · 3 weeks from kickoff to the AI owning the tier-one queue.
The problem
The AI wasn't connected to the company's actual knowledge. It answered from a model's general memory, not from the current docs, the real policy pages, or the thousand resolved tickets that already contained the right answers. And there was no measurement — quality was debated by anecdote in Slack, which meant nobody could say whether any change made things better or worse.
The engagement
Forge connected the AI to the truth: documentation, help center, policy pages, and a year of resolved tickets, kept in sync automatically as content changes. Every answer now cites the source it came from — visible to the customer, checkable by the team.
Just as important, quality became a number. An evaluation suite built from real customer questions scores the system continuously; any change that would degrade accuracy is caught before customers ever see it, not after.
The results
The measurement changed the conversation. We stopped debating whether an answer felt right and started reading the score. — Head of Support
Where it went next
With trust re-established, the second phase raised the ceiling: resolved conversations now feed the evaluation suite automatically, and the AI has graduated from answering questions to drafting full ticket replies for one-click human approval — the volume of automation, with the accountability of a person.
Golam Mostafa leads the engagements behind these results. Planning something similar? Talk it through first.