AigeoRadar

AigeoRadar · AI Context Benchmark

AI Context Benchmark

Question → sufficient layer(s) → retrieved slice → measured cost → evidence

Ankara Anahtar Teslim Ofis Tadilatı

https://ankaraanahtarteslim.com.tr/ankara-anahtar-teslim-ofis-tadilati/

9 questions Answer LLM: gpt-4o-mini ai_context_benchmark_v1

Of 9 questions: lowest measured retrieval cost among sufficient layers — HTML 1 · Schema 0 · AIPM 6 · multi-layer 0 · unanswered 2. Descriptive counts only; no overall ranking.

Of 9 questions: lowest measured retrieval cost among sufficient layers — HTML 1 · Schema 0 · AIPM 6 · multi-layer 0 · unanswered 2. Descriptive counts only; no overall ranking. Each question shows: sufficient layer(s) → retrieved slice → measured input tokens → evidence. Readers interpret; the report does not rank formats. Semantic Redundancy 56% — the manifest repeats the same ideas across fields. This inflates structural size without helping per-question retrieval.

Gold mode · Mixed independent + AIPM consistency

Most Understanding/Retrieval gold comes from HTML. Some questions remain AIPM self-consistency checks (tagged) and are excluded from Understanding when possible. Matching a machine card against its own fields is not evidence that that layer outperforms HTML.

Execution provenance

Methodology
AI Context Benchmark Methodology v1.0
Planner Spec
v1.0
Execution Protocol
v1.0
Question pack
universal_v1
Score version
aipm_benchmark_score_v10
Engine
aipm_benchmark_v4
Answer LLM
openai / gpt-4o-mini (single provider this lab)
Model snapshot
2026-07
Run id
ff29d5f0-364c-48db-bc3a-570643a72d6f

Question outcomes

Per question: which layers were sufficient, and what was the lowest measured retrieval cost among them. No overall winner.

1

Lowest cost: HTML

0

Lowest cost: Schema

6

Lowest cost: AIPM

0

Multi-layer

2

Unanswered

Manifest design

Semantic Redundancy · 56%

Semantic Redundancy 56% — the manifest repeats the same ideas across fields. This inflates structural size without helping per-question retrieval.

  • Merge or differentiate `purpose` and `abstract` (100% overlap).
Field A Field B Overlap
purpose abstract 100%
purpose keyFacts 47%
abstract keyFacts 47%
abstract title 28%

Structural size (secondary)

Structural size is secondary. A larger AIPM is not a failure if per-question retrieval stays tiny — check Semantic Redundancy instead.

HTML · 12,020 chars
Schema · 191 chars
AIPM · 12,143 chars

Routing helper (secondary)

Illustrative card-first vs HTML-always simulation — prefer per-question measured cost above. Not a ranking.

Card-first retrieval would use ~8% fewer tokens than HTML-always on this pack (7 answered from machine card, 2 escalated to HTML). Descriptive only.

Orientation pack

8 questions · all layers scored

  • HTML 3/8
  • Schema 3/8
  • AIPM 6/8

Depth pack

0 questions · all layers scored

  • HTML 0/0
  • Schema 0/0
  • AIPM 0/0

Full-context pack metrics (secondary)

These measure the whole file fed to the model this run — not the minimum slice needed per question.

HTML

3/9 matched

37,899 full-pack tokens

Schema

3/9 matched

1,489 full-pack tokens

AIPM

7/9 matched

34,794 full-pack tokens

Full-pack resource table (secondary)

Whole-file context fed this run. Prefer Minimal Retrieval Cost on each question.

Metric HTML Schema AIPM
Coverage (matched) 3/9 3/9 7/9
Context size 12,020 chars 191 chars 12,143 chars
Total tokens 37,899 1,489 34,794
Tokens / correct answer 12,633 496 4,971
Est. cost / correct answer $0.001922 $0.000088 $0.000756
Matched per 1k tokens 0.079 2.015 0.201
Median latency 1,550 ms 1,963 ms 1,807 ms
Est. cost (USD) $0.00577 $0.00026 $0.00529

Six independent scores

Answer Efficiency is the primary cost lens. Accuracy axes remain for research — no combined total or winner.

Answer Efficiency

Matched answers per 1k tokens (and cost per match). The primary efficiency axis — not raw accuracy.

  • HTML 3.9
  • Schema 100
  • AIPM 10

Understanding

Can this layer convey what the page is about — using independent HTML gold?

  • HTML 20
  • Schema 40
  • AIPM 60

Retrieval

Can this layer surface shared facts (location, contact, hours, pricing, FAQ, CTA)?

  • HTML
  • Schema
  • AIPM

Evidence

Answer quality vs independent gold (score strength). Partial credit counts; UNKNOWN scores zero unless gold is UNKNOWN.

  • HTML 46.2
  • Schema 35.9
  • AIPM 69.3

Metadata

Language, page kind, and freshness from page signals.

  • HTML 66.7
  • Schema 33.3
  • AIPM 100

Compression

Information delivered per token and context size. Higher means more matched answers for less context cost.

  • HTML 3.9
  • Schema 100
  • AIPM 10

Coverage map

AIPM matched 7 question(s) (alone on Q3, Q5, Q7); HTML matched 3. HTML/Schema (or a gap) still needed on Q1, Q2. No layer matched gold on Q1, Q2. Read this as complementary coverage — AIPM orients agents cheaply; HTML supplies depth when the sidecar cannot. HTML 3/9 matched (≈12,633 tok/match). Schema 3/9 matched (≈496 tok/match). AIPM 7/9 matched (≈4,971 tok/match). Figures are descriptive per layer — not a ranking. See per-question chains for sufficient layers, slice size, and measured cost.

AIPM matched

Q3, Q4, Q5, Q6, Q7, Q8, Q9

Alone: Q3, Q5, Q7

AIPM insufficient

Q1, Q2

HTML/Schema needed or all layers missed

Unanswered by all

Q1, Q2

HTML layer

Visible page text after stripping AIPM sidecars and JSON-LD. Measures what prose alone can answer.

Context fed: 12,020 chars

Tokens: 37,899 · median 1,550 ms

Matched this pack: 3/9

Stronger on

Weaker on

Understanding (20) · Compression (3.9) · Answer Efficiency (3.9)

Schema layer

JSON-LD structured data with minimal page chrome. Measures what schema markup can answer.

Context fed: 191 chars

Tokens: 1,489 · median 1,963 ms

Matched this pack: 3/9

Stronger on

Compression (100) · Answer Efficiency (100)

Weaker on

Metadata (33.3) · Evidence (35.9)

AIPM layer

AI Page Manifest (.ai.json) only. Measures what the machine layer can answer without HTML.

Context fed: 12,143 chars

Tokens: 34,794 · median 1,807 ms

Matched this pack: 7/9

Stronger on

Metadata (100)

Weaker on

Compression (10) · Answer Efficiency (10)

When to use which layer

AIPM complements HTML — it does not replace full-page prose.

Scenario Recommended Why
Fast orientation (title, purpose, brand, intent) Compare machine card → HTML fallback on this run Machine-card pack: 34,794 tok · $0.00529. Routing sim saved ~8% tokens vs HTML-always.
Deep content / research (prose facts, process detail) HTML (with optional machine orientation) HTML pack: 37,899 tok · $0.00577. Use when depth needs body prose.
Structured entity pulls (org, location, typed fields) Schema.org Schema pack: 1,489 tok · $0.00026. Dense JSON-LD tends to score well here.

Findings

  • Of 9 questions: lowest measured retrieval cost among sufficient layers — HTML 1 · Schema 0 · AIPM 6 · multi-layer 0 · unanswered 2. Descriptive counts only; no overall ranking.
  • Semantic Redundancy 56% — the manifest repeats the same ideas across fields. This inflates structural size without helping per-question retrieval.
  • Manifest design: Merge or differentiate `purpose` and `abstract` (100% overlap).
  • Structural size is secondary. A larger AIPM is not a failure if per-question retrieval stays tiny — check Semantic Redundancy instead.
  • AIPM matched 7 question(s) (alone on Q3, Q5, Q7); HTML matched 3. HTML/Schema (or a gap) still needed on Q1, Q2. No layer matched gold on Q1, Q2. Read this as complementary coverage — AIPM orients agents cheaply; HTML supplies depth when the sidecar cannot. HTML 3/9 matched (≈12,633 tok/match). Schema 3/9 matched (≈496 tok/match). AIPM 7/9 matched (≈4,971 tok/match). Figures are descriptive per layer — not a ranking. See per-question chains for sufficient layers, slice size, and measured cost.

Question-by-question layer analysis

Which layer(s) could answer; which need more or different context; minimum context fed this run.

Q1 · Understanding · orientation

Summarize the page in one sentence.

Gold: Amaç; Ankara anahtar teslim ofis tadilatı projelerinde tasarım, planlama ve uygulama süreçlerini tek elden yöneterek işletmelere hızlı, düzenli ve profesyonel çalışma ortamları kazandırmaktır.

Gold source: abstract · html_independent

No layer provided a sufficient answer from its context alone.

Estimated retrieval cost

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lead_paragraphs est no
Schema No JSON-LD on page. est no
AIPM abstract est no

Root cause · Layer miss

Improve orientation fields; do not paste full HTML into the sidecar.

HTML

insufficient

Ankara Anahtar Teslim Ofis Tadilatı offers comprehensive renovation and decoration services for office spaces, focusing on modern design, efficiency, and employee comfort since 2005.

score 24 · 12,020 chars context · 3,774 in-tokens

HTML did not answer from 12020 chars of context — additional or different layer context needed.

Schema

insufficient

The page is about the services offered by Ankara Anahtar Teslim related to office renovations.

score 21 · 191 chars context · 142 in-tokens

Schema did not answer from 191 chars of context — additional or different layer context needed.

AIPM

insufficient

Ankara anahtar teslim ofis tadilatı hizmetleri, eskiyen, verimsiz veya modern çalışma ihtiyaçlarına uygun olmayan ofis alanlarının baştan sona yenilenmesini kapsamaktadır.

score 30 · 12,143 chars context · 3,465 in-tokens

AIPM did not answer from 12143 chars of context — additional or different layer context needed.

Q2 · Understanding · orientation

What is the purpose of this page?

Gold: Amaç; Ankara anahtar teslim ofis tadilatı projelerinde tasarım, planlama ve uygulama süreçlerini tek elden yöneterek işletmelere hızlı, düzenli ve profesyonel çalışma ortamları kazandırmaktır.

Gold source: purpose · html_independent

No layer provided a sufficient answer from its context alone.

Estimated retrieval cost

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lead_paragraphs est no
Schema No JSON-LD on page. est no
AIPM purpose est no

Root cause · Layer miss

Improve orientation fields; do not paste full HTML into the sidecar.

HTML

insufficient

The purpose of this page is to provide information about Ankara Anahtar Teslim Ofis Tadilatı services, including details on the renovation and decoration of office spaces.

score 30 · 12,020 chars context · 3,772 in-tokens

HTML did not answer from 12020 chars of context — additional or different layer context needed.

Schema

insufficient

UNKNOWN.

score 0 · 191 chars context · 140 in-tokens

Schema did not answer from 191 chars of context — additional or different layer context needed.

AIPM

insufficient

The purpose of this page is to provide information about Ankara anahtar teslim ofis tadilatı hizmetleri, which involves the complete renovation of outdated or inefficient office spaces.

score 29 · 12,143 chars context · 3,463 in-tokens

AIPM did not answer from 12143 chars of context — additional or different layer context needed.

Q3 · Understanding · orientation

Who is the publisher or brand?

Gold: Ankara Anahtar Teslim

Gold source: publisher.name · html_independent

Sufficient: AIPM · Needs more / other context: HTML, Schema

Estimated retrieval cost · lowest cost AIPM · 6 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.title_brand est no
Schema No JSON-LD on page. est no
AIPM publisher 6 est yes

HTML

insufficient

ÖZ BALIKÇILAR İNŞAAT is the publisher or brand.

score 0 · 12,020 chars context · 3,771 in-tokens

HTML did not answer from 12020 chars of context — additional or different layer context needed.

Schema

insufficient

UNKNOWN.

score 0 · 191 chars context · 139 in-tokens

Schema did not answer from 191 chars of context — additional or different layer context needed.

AIPM

sufficient

The publisher is Ankara Anahtar Teslim.

score 88 · 12,143 chars context · 3,462 in-tokens

AIPM answered using 12143 chars of layer context (minimum fed this run).

Q4 · Metadata · orientation

What language is this page in?

Gold: tr

Gold source: inLanguage · html_independent

Sufficient: HTML, AIPM · Needs more / other context: Schema

Estimated retrieval cost · lowest cost AIPM · 1 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lang 2 est yes
Schema No JSON-LD on page. est no
AIPM inLanguage 1 est yes

HTML

sufficient

The page is in Turkish.

score 88 · 12,020 chars context · 3,771 in-tokens

HTML answered using 12020 chars of layer context (minimum fed this run).

Schema

insufficient

UNKNOWN.

score 0 · 191 chars context · 139 in-tokens

Schema did not answer from 191 chars of context — additional or different layer context needed.

AIPM

sufficient

The page is in Turkish (tr).

score 88 · 12,143 chars context · 3,462 in-tokens

AIPM answered using 12143 chars of layer context (minimum fed this run).

Q5 · Metadata · orientation

What page type is declared in the machine layer?

Gold: blog

Gold source: pageType · aipm_consistency

Sufficient: AIPM

Estimated retrieval cost · lowest cost AIPM · 1 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lead est no
Schema No JSON-LD on page. est no
AIPM pageType 1 est yes

HTML

skipped

UNKNOWN.

score 0 · 0 chars context · 0 in-tokens

Skipped (AIPM-only).

Schema

skipped

UNKNOWN.

score 0 · 0 chars context · 0 in-tokens

Skipped (AIPM-only).

AIPM

sufficient

The page type is "blog."

score 88 · 12,143 chars context · 3,461 in-tokens

AIPM answered using 12143 chars of layer context (minimum fed this run).

Q6 · Understanding · orientation

What is the page title?

Gold: Ankara Anahtar Teslim Ofis Tadilatı – Ankara Anahtar Teslim

Gold source: title · html_independent

Sufficient: Schema, AIPM · Needs more / other context: HTML

Estimated retrieval cost · lowest cost AIPM · 9 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.title est no
Schema No JSON-LD on page. est yes
AIPM title 9 est yes

HTML

insufficient

UNKNOWN

score 0 · 12,020 chars context · 3,770 in-tokens

HTML did not answer from 12020 chars of context — additional or different layer context needed.

Schema

sufficient

The page title is "Ankara Anahtar Teslim Ofis Tadilatı – Ankara Anahtar Teslim."

score 93 · 191 chars context · 138 in-tokens

Schema answered using 191 chars of layer context (minimum fed this run).

AIPM

sufficient

The page title is "Ankara Anahtar Teslim Ofis Tadilatı."

score 91 · 12,143 chars context · 3,461 in-tokens

AIPM answered using 12143 chars of layer context (minimum fed this run).

Q7 · Understanding · blog

What does the short summary say?

Gold: Ankara anahtar teslim ofis tadilatı hizmetleri; eskiyen, verimsiz veya modern çalışma ihtiyaçlarına uygun olmayan ofis alanlarının baştan sona yenile…

Gold source: summary150 · unknown

Sufficient: AIPM · Partial: HTML · Needs more / other context: Schema

Estimated retrieval cost · lowest cost AIPM · 128 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lead est no
Schema No JSON-LD on page. est no
AIPM orientation_bundle 128 est yes

HTML

partial

Öz Balıkçılar İnşaat, 2005’ten beri Ankara’da tadilat ve dekorasyon alanında hizmet vermektedir.

score 55 · 12,020 chars context · 3,771 in-tokens

HTML was only partially aligned with gold from 12020 chars of context.

Schema

insufficient

UNKNOWN.

score 0 · 191 chars context · 139 in-tokens

Schema did not answer from 191 chars of context — additional or different layer context needed.

AIPM

sufficient

Ankara anahtar teslim ofis tadilatı hizmetleri; e…

score 80 · 12,143 chars context · 3,462 in-tokens

AIPM answered using 12143 chars of layer context (minimum fed this run).

Q8 · Understanding · orientation

What is the primary topic of this page?

Gold: Ankara Anahtar Teslim Ofis Tadilatı

Gold source: primaryTopic · html_independent

Sufficient: HTML, Schema, AIPM

Estimated retrieval cost · lowest cost HTML · 4 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.h1|title 4 est yes
Schema No JSON-LD on page. est yes
AIPM primaryTopic 6 est yes

HTML

sufficient

The primary topic of this page is Ankara Anahtar Teslim Ofis Tadilatı services.

score 84 · 12,020 chars context · 3,773 in-tokens

HTML answered using 12020 chars of layer context (minimum fed this run).

Schema

sufficient

The primary topic of this page is "Ankara Anahtar Teslim Ofis Tadilatı."

score 85 · 191 chars context · 141 in-tokens

Schema answered using 191 chars of layer context (minimum fed this run).

AIPM

sufficient

The primary topic of this page is "Ankara Anahtar Teslim."

score 60 · 12,143 chars context · 3,464 in-tokens

AIPM answered using 12143 chars of layer context (minimum fed this run).

Q9 · Metadata · orientation

What is the content intent?

Gold: informational

Gold source: contentIntent · html_independent

Sufficient: HTML, Schema, AIPM

Estimated retrieval cost · lowest cost AIPM · 4 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.prose_signals 100 est yes
Schema No JSON-LD on page. est yes
AIPM contentIntent 4 est yes

HTML

sufficient

The content intent is to provide information about Ankara Anahtar Teslim Ofis Tadilatı services, including details on the services offered, processes, and benefits of office renovation.

score 88 · 12,020 chars context · 3,770 in-tokens

HTML answered using 12020 chars of layer context (minimum fed this run).

Schema

sufficient

The content intent is to provide information about the services related to office renovation in Ankara.

score 88 · 191 chars context · 138 in-tokens

Schema answered using 191 chars of layer context (minimum fed this run).

AIPM

sufficient

The content intent is informational.

score 88 · 12,143 chars context · 3,461 in-tokens

AIPM answered using 12143 chars of layer context (minimum fed this run).

Methodology

  • AI Context Benchmark: Planner → Slice → LLM with measured API input tokens.
  • Reports describe sufficient layers, slice size, cost, and evidence — they do not declare a winning format.
  • Layers under test today: HTML, Schema.org, AIPM (extensible to RSS, Markdown, PDF, …).
  • Engine aipm_benchmark_v1 · Tue, Jul 28, 2026 4:10 PM · aipm_benchmark_score_v10

Share

Per-question context chain across layers.