Case study

GridResolve AI case study

GridResolve AI is a multi-agent prototype for resolving utility billing disputes, built on Microsoft Foundry. It was submitted to the Microsoft Agent-a-thon at the Level 3 Architect tier. It is an independent project, not a Brandsap product.

The problem

Billing disputes at a utility touch meter data, tariffs, account history and regulation. An AI assistant that drafts a confident but unsupported answer to a customer is worse than no assistant at all.

What I built

A set of cooperating agents that gather evidence, check it against tariff and compliance rules, and draft a resolution, plus a release gate that blocks any answer whose claims are not backed by recorded evidence. A static control center shows the workflow.

My role

Sole designer and builder: agent design, evaluation criteria, release gate, documentation and the submission.

Architecture

  • Nine agents on Microsoft Foundry using gpt-5-mini, each with a narrow responsibility.
  • A claim ledger that records which piece of evidence supports each statement in a draft.
  • A fail-closed release gate: if any claim lacks support, the answer is held for a human instead of being sent.
  • Acceptance criteria written down before the build, so the evaluation could not be tuned after the fact.

Technologies

Microsoft Foundrygpt-5-miniMulti-agent orchestrationStatic HTML control center

How it works

A dispute enters the workflow, specialist agents collect and check the relevant facts, each claim in the proposed answer is linked to its evidence in the ledger, and the release gate decides whether the answer may be released or must go to a person.

Key engineering decisions

  • Fail closed rather than fail open: an unsupported claim stops the release.
  • Pre-registered acceptance criteria to keep the evaluation honest.
  • Synthetic data only, so no real customer information is used anywhere in the prototype.

Challenges

  • Keeping nine agents from repeating or contradicting each other while staying within a small model's context.
  • Making the release decision explainable to a reviewer, not just a pass or fail.

What I learned

For regulated workflows the gate matters more than the generator. A plain, auditable rule for when an answer is allowed out is what makes an LLM system usable.

Current status

Prototype using synthetic data. Submitted to the Microsoft Agent-a-thon (Level 3 Architect). No placement has been announced or claimed.

Links

More case studies: Capital Intelligence OS · Rankelo · Vaani · TalkBot · Bizia