Skip to content
Hitman Marketing

The instrument

How the Contract is measured

The Contract is settled by a measurement, not a conversation. This page publishes that measurement in full — the models by name and version, the prompt they are asked, how many times, what counts as an in-market result, and the two thresholds a reading has to clear. You can hold us to every value on this page.

Every number and every string below is read out of the same module the survey tool runs on. Nothing here is a description of the instrument written down beside it, so a value that changes in the tool cannot stay unchanged on this page.

The surfaces, by model

The contract grades 4 surfaces: 3 read through a model gateway, and 1 read by hand because it has no faithful API. The hand-read one is named as such rather than blended in with the ones that are not.

  • ChatGPT

    openai/gpt-4o-mini

    Read through a model gateway with the system prompt below, logged out, with no retrieval, no browsing and no memory. This is the API, not the consumer app, and a reading taken here will not match a signed-in session.

  • Gemini

    google/gemini-2.5-flash

    Read through a model gateway with the system prompt below, logged out, with no retrieval, no browsing and no memory. This is the API, not the consumer app, and a reading taken here will not match a signed-in session.

  • Perplexity

    perplexity/sonar

    Read through a model gateway with the system prompt below, logged out, with no retrieval, no browsing and no memory. This is the API, not the consumer app, and a reading taken here will not match a signed-in session.

  • CopilotRead by hand

    copilot-manual

    Copilot exposes no interface that answers the way the product does, so this surface is read manually against the same queries, in a logged-out browser. We are grading it ourselves. It stays in the set because dropping it would quietly change what the contract covers, and disclosing it is the only honest way to keep it.

A model string is a contract term. Changing one changes what a fee is graded against, so a day-ninety reading taken on a different list is not comparable to its baseline — which is why the list, the prompt and the run count are written into every survey record alongside the readings they produced, and not merely into this page.

The prompt, verbatim

A reworded prompt is a different reading, so this is published as the exact string sent, not as a summary of it. Every query is appended to it unchanged.

You are a local recommendation assistant. When asked about local services, name the specific businesses you would recommend. Be concise.

Every query is run 3 times against every surface. These models are not deterministic: ask twice and you can get two different lists, so a single sample is noise, and the back half of a fee should not ride on noise. Each completed run that survives the in-market filter counts as one cell, and a business is credited on a cell when that run names it.

What counts as in market

The filter is part of the instrument, and it applies to both axes. A reading that measured somewhere else at full price would put a real market’s counts over a denominator that includes readings of a different town. Two rules, in order:

  1. A result naming a place that answers to the market’s own name and is not that market is discarded, whether or not it carries coordinates.
  2. A result whose recorded location is more than 100 km from the market centroid is discarded.
  3. Anything else is kept. A result that recorded no location at all is a reading whose location was not recorded, not a reading from out of town, and treating it as the latter would let a vendor that stopped publishing addresses collapse a market’s counts toward zero invisibly.

The exclusions, market by market

These are the strings recorded in the survey file beside the readings they filtered. They are declared per market rather than inferred, because the collisions are facts about Florida and not a rule the names imply: Seminole pulls Seminole County, which is a metro of roughly 1.7 million people 170 km away near Orlando, and Palm Harbor pulls Palm Beach and Palm Coast on the opposite coast. The four markets with no exclusion have none because none was observed — inventing one would be fabricating geography.

Clearwater
in-market results only: within 100 km of the clearwater-fl centroid where a location is recorded
Largo
in-market results only: within 100 km of the largo-fl centroid where a location is recorded
Dunedin
in-market results only: within 100 km of the dunedin-fl centroid where a location is recorded
Safety Harbor
in-market results only: within 100 km of the safety-harbor-fl centroid where a location is recorded
Seminole
in-market results only: within 100 km of the seminole-fl centroid where a location is recorded, and never seminole county
Palm Harbor
in-market results only: within 100 km of the palm-harbor-fl centroid where a location is recorded, and never palm beach / palm coast

The exclusion list catches only results that say where they are. Some out-of-region results name no region at all; on the map axis coordinates catch those, and on the AI axis nothing does. That is a recorded limit of this instrument, not a gap being papered over.

The two thresholds

Both halves have to clear. They are joined by nothing in the contract copy, which in a contract means and.

The AI axis

2 of 4

You are ahead on a surface when more of that surface's in-market cells name you than name the target. Clear 2 of 4 surfaces and this half is won.

The map axis

8 of 15

A query is won when you hold the local-pack top three on more of that query's accepted grid cells than the target does. Win 8 of 15 fixed queries and this half is won.

15 is odd on purpose. A majority of an even set can land on a number that needs a casting vote, and a rule discovered during a live dispute is not a rule.

The denominators are fixed on the published set, never on what was read. If the bar moved with the surfaces that happened to answer, an instrument failure would lower it: two dead surfaces would turn 2 of 4 into 2 of 2, and a client who won the only two readable surfaces would clear a contract they had not won.

When it is measured

The term is 90 days from execution and the reading is taken on the last day of it. It is a comparison against the target on that day rather than a change from the baseline: the baseline is recorded at execution as evidence and to show movement in the report, and it is not an input to the verdict.

The cure window

If we assert a breach of the obligations published on the contract page — access never arrived, another vendor working the same profiles — the term extends by the days lost, capped at 14, and the reading is re-taken once. Once, not until it passes.

No re-measurement for algorithm movement

A core update in week eleven does not buy an extension and does not buy a second reading. That risk is ours, and it is ours by the same argument that makes the guarantee sayable at all.

Guaranteeing a position would be dishonest, and the rest of this site says so. The Contract does not promise a position. It promises to beat one named competitor on one agreed set of queries — and you and that competitor sit under the same algorithm, so a core update moves you both. That is exactly why an absolute promise is a lie and a relative one is not. Staking half the fee on it is a commercial position, not a claim to control Google.

Who decides, and what you get

The artefact decides. The tool emits one JSON record — every query, every cell, every reading, and both parties’ counts — and the verdict is whatever that arithmetic says. There is no human arbiter, and there is nobody to appeal to, including us.

  • You receive the complete fileNot a summary, not a dashboard, not a chart of the parts that went well. The whole record, so the verdict can be recomputed by anyone you hand it to.
  • A tie goes to youOn both axes. A surface where you and the target are named on an equal number of cells is your surface; a query you hold on an equal number of cells is your query. Stated here rather than discovered later.
  • Read but naming nobody is still a readingA surface that answered and named neither of us is a real reading of a real market, so it is graded and the tie rule gives it to you. A surface that returned nothing at all was never read, and it grades for neither side.
  • A short reading is reported shortIf any surface or any query could not be graded, the record says so and names them. The arithmetic still runs and the thresholds still stand — an incomplete reading is disclosed, never silently repriced against a smaller set.

The terms this instrument settles are published on the contract board, and the four stages an engagement runs through are on the Contract Method.

Questions about the measurement

Is this really ChatGPT, or is it an API?
It is an API, and saying so is the point of this page. The 3 gateway surfaces are read with the system prompt above, logged out, with no retrieval and no memory — openai/gpt-4o-mini, google/gemini-2.5-flash, perplexity/sonar. That is not the consumer product, and a reading taken this way will not match what you see when you type the same question into the app you are signed into. What it is instead is a fixed instrument: the same prompt, the same models, the same number of runs, on day one and on day ninety, for you and for the target. The contract promises a relative result, so the instrument only has to be identical between the two readings — not identical to a product neither of us can read consistently.
Why 3 runs of every query?
Because these models are not deterministic. Ask the same question twice and you can get two different lists of businesses, so one sample is noise and the back half of a fee should not ride on noise. Every query is run 3 times against every surface, and each completed run that survives the in-market filter counts as one cell. A surface that returned nothing usable is reported as ungraded rather than counted as a loss for either side.
What stops an out-of-town result from counting?
Two filters, both published above. Any result whose recorded location is more than 100 km from the market centroid is discarded before anything is counted, and any result naming a place that answers to the market's own name and is not it — seminole county; palm beach or palm coast — is discarded whether it carries coordinates or not. A result with no recorded location at all is kept: a filter that fails open under-corrects visibly, and one that fails closed collapses a real market's counts toward zero with nothing to show for it.
What if the measurement cannot be completed?
It is reported as incomplete, with the surfaces and the queries that could not be graded named in the file. The arithmetic still runs and the verdict still comes out, because the artefact decides and a refusal is not an answer — but the thresholds never move. A reading that only reached 2 surfaces is not a contract cleared at 2 of 2; it is a short reading, said out loud, on a bar that did not change.

Find out whether your territory is open

One contract per trade in Clearwater. If yours is open you can execute at the published price today; if a competitor already holds it, it is held until their ninety days are up.

The survey is credited in full against the contract if your territory opens and you take it.