gudteck

Gudteck Research · Pilot study

What AI tells New Zealand about its biggest real-estate brands

We asked the questions real buyers and vendors ask, scored what the assistant said about the five largest brands — and checked every claim it made.

In brief. 59 realistic buyer and vendor questions — 11 multi-turn conversations plus 10 one-shot benchmarks — put to a major AI assistant with live web search. Even the most visible brand appeared in barely a third of high-intent answer moments. No firm publishes the commercial facts AI needs, so third parties author them: the assistant described one firm's fee model three different ways in three conversations. Every conversation was logged; every claim about a firm was verdicted against primary sources.

The scoreboard

Five brands, anonymised here as Firms A–E. Each firm can request its own identified report.

FirmVendor recommendation /100Buyer visibility /100Accuracy /100 (claims checked)Unanchored claimsAI reachability /10
Firm A35.225.683.3 (3)34
Firm B29.621.2— (0)27
Firm C22.810.9100.0 (1)39
Firm D21.916.759.1 (11)39
Firm E6.20.770.0 (5)45

Scores are the weighted share of high-intent answer moments captured. Accuracy samples are small at pilot scale — counts in brackets. "Unanchored claims" are statements AI made that nobody can verify, because the firm publishes nothing to check the claim against.

Even the leader is in the room for barely a third of the conversations that decide who gets hired.

What we found

Nobody owns their own numbers.

None of the five publishes its fees, so the assistant answered every fee question anyway — from calculator sites, referral blogs and stale marketing — describing Firm E's fees three different ways, and calling Firm B the most expensive option using a number assembled entirely from third parties. Firm B's accuracy score is blank for the same reason: nothing it said about their numbers could be checked against anything the firm publishes.

Reachability is not authority.

The most visible brand (A) had the worst technical plumbing — its firewall rejected live AI retrieval, and its score was carried almost entirely by franchise sites and portal listings, assets head office doesn't control. The best-equipped site (D, 9/10) carried the least accurate story: its stated size was years stale and an entire regional expansion was invisible in every relevant answer.

The advice moments are unclaimed.

Questions about methods of sale, market timing, legal due diligence and property risk went almost entirely to portals, banks, law firms and government guides. Exactly one piece of agency content framed an answer in the whole study — a single franchise office's market update.

Aggregators are capturing the choosing moment.

Asked "which agency should I use?", the assistant repeatedly routed vendors to referral and review platforms — intermediaries that charge agencies for the lead — including one paid-referral service that took the recommendation outright. Some of those platforms themselves block AI crawlers.

The rails are open. The ground is unclaimed.

Trade Me explicitly allowlists AI crawlers on its listings and the major portals are readable — inventory reaches AI-assisted buyers regardless of what agencies do. The contested ground is brand, recommendation, facts and advice, and none of the five firms had an llms.txt. At the time of fieldwork, the category's machine-readable layer was effectively empty.

A note on our own reviewer

While preparing this publication, we asked an AI assistant to critically assess our own company. Its review was sharp and largely fair — and it confidently named a competitor that, as far as we can establish, does not exist. Even the audit of the auditor contained an unanchored claim. That is the category, in one sentence.

Method, in the open

Short version below. The instrument, rubric and per-turn observations exist in full and travel with every identified report.

Instrument
59 scored queries — 32 buyer-side and 17 vendor-side turns across 11 persona-based, multi-turn sessions, plus 10 one-shot benchmarks; context carries between turns like a real conversation.
Subjects
New Zealand's five largest residential brands by market presence, never named to the assistant except where a question naturally names them — and those turns are excluded from visibility scoring.
Execution
Fresh, memory-free sessions on a major AI assistant with live web search; the operators running sessions were blind to the target-firm list; single run per session — a directional baseline, not the full protocol.
Scoring
A firm counts as present only in the rendered answer or its citations; 0–3 points per appearance, mapped by question type and weighted by purchase intent.
Claims
Every factual claim verdicted: correct, partial, incorrect — or unverifiable, meaning the firm publishes nothing to check it against — excluded from the accuracy scores and reported separately as unanchored claims.
Reachability
Each firm's own web estate audited against seven technical checks — crawler policy, live AI retrieval, search-index presence, server-rendered content, structured data, sitemaps, llms.txt — scored /10.
Reproducibility
All scores recompute by script from per-turn observations. Our first hand-summed aggregates contained four errors; the published figures are the recomputed ones, and the correction stays in the record — it is why the pipeline exists.

Limitations, plainly: one AI platform, one run per session (assistants are stochastic — the full protocol runs three repetitions across ChatGPT, Google AI Mode and Perplexity) · scored by the study's designers against a published rubric, no independent second scorer yet · clean-account lab conditions with scripted phrasing, not field observation (a naturalistic-phrasing arm exists in the v1.2 instrument to measure that gap) · per-firm accuracy samples are small (1–11 claims) · retrieval ran against a predominantly US-weighted search index.

What's next

Tier-1 fieldwork: the same instrument on ChatGPT, Google AI Mode and Perplexity, three repetitions per session inside a 14-day window, with per-platform scores and a deepened claims register. Firms covered by this study can request their identified scorecard and evidence file — every transcript, every claim, every verdict — as a findings-and-recommendations report, NZ$2,000 +GST, with implementation consulting from there.

Cite as: Gudteck (2026). AI Visibility in New Zealand Residential Real Estate — Pilot Study. Published 13 August 2026. https://gudteck.com/research/nz-real-estate-pilot/

Want this run on your market?

Any industry where customers ask AI before they buy. Same instrument, your questions, your competitors.

Start the conversation