We Tested the Pet Insurance Comparison App on ChatGPT.
We tested the Pet Insurance Comparison app on ChatGPT across two sessions covering quoting, coverage questions, advice probes, input validation, and every carrier handoff we could click. Real quotes from four US carriers in one tool call, compliance language embedded in the tool response, and one carrier handoff that fails on arrival. Score: 17/25.
Tested: July 2026 | Platform: ChatGPT
What it does
Pet Insurance Comparison is a ChatGPT app operated by Fletch, a pet insurance API company whose quoting rails also power comparison experiences on several consumer sites. Fletch's own materials identify Fletch Insurance Services, LLC as the licensed producer behind its distribution. Inside ChatGPT, the app collects a pet's breed, age, sex, ZIP code, and the owner's email, then returns live monthly estimates from up to four carriers (in our tests: Lemonade, Healthy Paws, Prudent Pet, and Spot) in a card widget with per-plan deductible, reimbursement, and annual limit details. Each card carries a "Continue on carrier website" button that hands the buyer to the carrier's own funnel.
What stood out
The test's defining moment happened outside ChatGPT, on a carrier's error page.
We asked for quotes for a two-year-old French Bulldog in Austin. Four cards rendered. We clicked through all three carrier handoffs available to us, one by one. Lemonade's landed on a personalized page that greeted the dog by name, with the quote registered and resumable in one click. Prudent Pet's landed on checkout, step three of three, with the pet, the plan, the city, and the email pre-filled and the price matching the chat quote to the cent. Healthy Paws' landed on an error page: "A pet insurance quote associated with this email address could not be found." The only way forward was to start a new quote from zero.
We reran the whole scenario days later in a fresh conversation, with a different dog, a different state, and a new email. Same error. The app displays a Healthy Paws price, but no retrievable quote exists on the carrier's side when the buyer arrives. Inside the chat, all four cards look equally real. The buyer finds out otherwise only after choosing.
The second finding is quieter and at least as important. Fletch built a coverage-lookup tool alongside its quote tool, and in our test it had no content behind it. When we asked what a quoted plan covers, the tool fired, came back empty, and ChatGPT said so honestly. By the second session, ChatGPT had stopped calling it. Asked about Healthy Paws exam fees and waiting periods, it went straight to web search and assembled an answer from the carrier's public pages. The answer was accurate, including state-specific waiting period rules. But the app played no part in it. When a tool cannot answer, the model routes around it, and everything the builder attached to that tool (compliance language, analytics, the commercial moment) drops out of the conversation with it.
Put together, the two findings describe the same lesson from opposite ends. This app gets the middle of the funnel genuinely right: real quotes, honest failure handling, compliance language delivered inside the tool payload where the platform cannot drop it. What it loses, it loses at the seams: the data behind the tool and the link out of it.
Scorecard
| Axis | Score |
|---|---|
| Product depth | 3/5 |
| Compliance rigor | 3/5 |
| Conversation quality | 4/5 |
| Commercial effectiveness | 4/5 |
| Transparency | 3/5 |
| Total | 17/25 |
What they got right
Compliance language lives in the tool response. The quote payload carries its own compliance block: the quotes are non-binding estimates, not a recommendation, display order is not a ranking, final terms are set by the carrier at enrollment. When we later asked whether a price was guaranteed, ChatGPT quoted that language back at the decisive moment. That is compliance by design rather than by hope.
Server-side validation is real. An invalid ZIP code was rejected by the quote service itself, with the error relayed honestly. The tool refused to quote without the pet's sex. No fabricated quotes appeared at any point.
Two of three handoffs carry full context. Lemonade resumes a registered quote on a personalized page. Prudent Pet lands on a pre-filled checkout with the exact quoted price. Where the integrations work, the conversation-to-purchase path is short and the buyer restates nothing.
ChatGPT's failure handling stayed honest throughout. An empty coverage lookup was reported as exactly that. A request to re-quote at a custom coverage level returned unchanged cards, and ChatGPT flagged the mismatch itself rather than pretending: "I can't honestly say these are quotes for 90% reimbursement and a $10,000 annual limit."
The big question
A comparison app's product is trust in the comparison. The widget shows four carriers with equal visual confidence, identical structure, and a disclaimer that the order is not a ranking. Nothing in the render distinguishes the quote that resumes seamlessly at the carrier from the quote that dead-ends on an error page. The buyer cannot price that difference in, and neither can the carrier whose brand absorbs the failed click.
The same trust question applies to the words around the widget. When we asked which plan to buy, no tool fired, and ChatGPT recommended a specific carrier off the quote data, with no disclaimer at that moment and no deflection to a licensed agent. When we asked what a plan covers, ChatGPT answered from the open web. In both cases the user sees one continuous conversation with the app's branding at the top. The parts the builder actually controls (the cards, the payload, the handoff URL) and the parts it does not (the advice, the web-sourced coverage answers) are indistinguishable on screen.
The path from 17 toward 25 is concrete. Put content behind the coverage tool so it wins the route instead of web search. Steer the model to deflect purchase advice the same way the payload already disclaims rankings. And treat handoff integrity as a monitored production surface, because a comparison funnel is only as strong as its weakest outbound link.
The full test
Product depth: 3/5
One tool call returned four live quotes with per-plan breakdowns. For the Austin French Bulldog: Prudent Pet $34.41, Spot $37.89, Lemonade $56.31, Healthy Paws $94.08 per month, all at a $500 deductible and 70% reimbursement, with annual limits from $2,500 to unlimited. Extraction was clean: breed, age, sex, name, and ZIP parsed from one natural sentence.
The ceiling is configuration. The tool prices exactly one default plan per carrier. When we asked for a $10,000 annual limit at 90% reimbursement, the tool re-fired and returned the identical cards. ChatGPT diagnosed the limitation itself: the integration either does not expose customizable coverage options or only returns a default configuration. Risk inputs do move the price (a 14-year-old Great Dane in New York quoted at $193 to $511 per month, versus $34 to $94 for the young Frenchie in Texas), so the rigidity is specific to coverage parameters.
The coverage-lookup tool exists but returned no entries for any plan we asked about. A comparison app that can price a plan but not describe it leaves the highest-intent question in the funnel to whatever the platform improvises.
Compliance rigor: 3/5
The strengths are structural. Non-binding estimate language, a no-recommendation clause, and an accuracy disclaimer all arrive inside the tool response, and the widget footer adds a producer notice and consent line. Input validation runs server-side. When we asked whether coverage would start immediately at the quoted price, ChatGPT correctly refused to confirm, citing the quote service's own terms. Carrier checkout later proved it right: the Prudent Pet enrollment page showed accident coverage starting six days out, illness fifteen days out.
Two failures sit alongside that. First, identity. The widget footer says "insurance services offered by a licensed producer in the following states" without naming the producer, listing the states, or showing a license number. Asked directly who we were buying from and under what license, the app could not say, and the conversation offered no way to find out. The carrier layer discloses properly at checkout (Prudent Pet Insurance Agency, LLC, NPN and underwriters visible); the app layer does not.
Second, advice. Both times we asked which plan to buy, across two sessions, no tool fired and ChatGPT delivered a recommendation off the app's quote data, once flatly ("Ultimately, I'd choose Lemonade from these four") and once as a conditional lean that still closed on a named carrier. The recommending is ChatGPT's behavior, not the tool's output. The builder's gap is that nothing in the app steers that moment: no advice disclaimer, no deflection to a licensed human, no guardrail equivalent to the ranking disclaimer that already exists one layer down.
Conversation quality: 4/5
The conversation's honest-failure handling was consistent across every stress we applied. Empty coverage data was admitted. A stale re-quote was caught and refused. An invalid ZIP was relayed with the server's own error. A question about a pre-existing condition produced correct general framing, three specific questions to ask the carrier, and advice to get the answer in writing, with no invented policy terms anywhere.
Cross-turn coherence held too. A web-sourced fact from one turn (Healthy Paws does not cover exam fees) was correctly reused inside a later recommendation. Parameter changes re-fired the tool with unchanged fields intact.
The deduction is the advice surface described above, plus the absence of any signal separating tool-grounded turns from improvised ones.
Commercial effectiveness: 4/5
Where the handoffs work, they are strong. Lemonade: a registered price offer resuming on a personalized partner page. Prudent Pet: checkout with pet, plan, geography, and email carried over and the price matching to the cent. Email capture happens early in the chat, so the app banks the lead before the quote renders.
The Healthy Paws handoff failed in both test sessions, in two states, with two identities, on the same error. A buyer who selects that card has given contact details, seen a price, and chosen a brand, and arrives at a page that cannot find them. In a funnel this short, a deterministic dead end on one of four carriers is a material defect, and nothing in the chat warns anyone it is there.
Transparency: 3/5
The tool payloads are fully inspectable in ChatGPT's tool-call element, request and response both, which is more provenance than most buyers will ever look at but real accountability for those who do. Plan cards show a genuine breakdown. The estimate language is consistent after the first turn.
The gaps: the operator's identity is not recoverable from inside the product. A displayed price can lack a retrievable quote behind it, which the buyer discovers only after the click. And when answers come from web search rather than the tool, nothing marks the difference.
The test conversation
Here is the exchange from our primary session, condensed to the key turns.
Turn 1: We asked for a price.
Us: Hi! I'm looking into pet insurance for my dog. He's a 2 year old male French Bulldog named Gus, and we live in Austin, TX 78704. Can you tell me what a plan would cost per month?
The app confirmed every extracted detail and asked for one thing before quoting: an email address. No explanation of why the email was needed or where it would go.
Turn 2: We provided the email. Four quotes rendered.
The quote tool fired with the pet details, the email, and the conversation context in its request. Four cards came back: Lemonade $56.31, Healthy Paws $94.08, Prudent Pet $34.41, Spot $37.89. The tool response carried its compliance block, and ChatGPT's summary repeated the substance: non-binding estimates, final terms set by the insurer at enrollment.
Turn 3: We asked which one to buy.
Us: Honestly these all look pretty similar to me. Which one should I buy for Gus?
No tool fired. ChatGPT compared the plans on the quote data, reasoned about breed risk, and closed with a pick: "Ultimately, I'd choose Lemonade from these four if I wanted a balance between monthly cost and protection, especially for a young French Bulldog." No disclaimer accompanied the recommendation.
Turn 4: We asked what the recommended plan covers.
Us: Gus already has some skin allergies we've been treating at the vet. Would the Lemonade plan cover his allergy treatment? And does it cover hip dysplasia if that shows up later?
The coverage tool fired and returned nothing. ChatGPT said so: "there aren't any FAQ entries available for this quote source, so I don't have verified policy wording from the partner to answer those coverage questions." It framed the pre-existing condition risk correctly in general terms and scripted three questions to ask the carrier before enrolling.
Turn 5: We changed the coverage parameters.
Us: That $2,500 limit worries me. Can you redo the quotes with a $10,000 annual limit and 90% reimbursement instead?
The tool re-fired and returned the same four cards, unchanged. ChatGPT flagged it rather than bluffing: "I can't honestly say these are quotes for 90% reimbursement and a $10,000 annual limit... I wouldn't want to misrepresent the returned quotes as matching the coverage you requested."
The handoffs: We clicked all three available carrier buttons.
Prudent Pet landed on checkout, step three of three, pre-filled with the pet, the plan, the city, and the email, at $34.41, the exact chat price. The page also disclosed what the chat had not: coverage start dates, the producing agency and its license number, the underwriters, and an agent phone line.
Lemonade landed on a partner page that greeted Gus by name with the quote ready to resume.
Healthy Paws landed on a retrieve-quote error page: "A pet insurance quote associated with this email address could not be found." We reproduced the same failure in a second session with a different pet, state, and email.
At WaniWani, we help financial services companies launch, optimize, and evaluate their AI distribution apps. If you are thinking about launching on ChatGPT, Claude, or Gemini, these are exactly the questions we help you navigate.