AigeoRadar

AigeoRadar · AI Context Benchmark

AI Context Benchmark

Question → sufficient layer(s) → retrieved slice → measured cost → evidence

Tij galvaniz

https://kapadokyagalvaniz.com.tr/tis-galvaniz/

10 questions Answer LLM: gpt-4o-mini ai_context_benchmark_v1

Of 10 questions: lowest measured retrieval cost among sufficient layers — HTML 3 · Schema 1 · AIPM 6 · multi-layer 0 · unanswered 0. Descriptive counts only; no overall ranking.

Of 10 questions: lowest measured retrieval cost among sufficient layers — HTML 3 · Schema 1 · AIPM 6 · multi-layer 0 · unanswered 0. Descriptive counts only; no overall ranking. Each question shows: sufficient layer(s) → retrieved slice → measured input tokens → evidence. Readers interpret; the report does not rank formats. Semantic Redundancy 66% — the manifest repeats the same ideas across fields. This inflates structural size without helping per-question retrieval.

Gold mode · Independent HTML gold

Gold answers are derived from the HTML page (title, h1, lang, prose) — not from AIPM fields. Understanding/Retrieval therefore measure layer capability, not AIPM self-consistency. Human-authored gold.json remains the gold standard for publication-grade claims.

Execution provenance

Methodology
AI Context Benchmark Methodology v1.0
Planner Spec
v1.0
Execution Protocol
v1.0
Question pack
universal_v1
Score version
aipm_benchmark_score_v10
Engine
aipm_benchmark_v4
Answer LLM
openai / gpt-4o-mini (single provider this lab)
Model snapshot
2026-07
Run id
9eb9d16e-cbae-4a7f-9186-d9f65d8ab277

Question outcomes

Per question: which layers were sufficient, and what was the lowest measured retrieval cost among them. No overall winner.

3

Lowest cost: HTML

1

Lowest cost: Schema

6

Lowest cost: AIPM

0

Multi-layer

0

Unanswered

Manifest design

Semantic Redundancy · 66%

Semantic Redundancy 66% — the manifest repeats the same ideas across fields. This inflates structural size without helping per-question retrieval.

  • Merge or differentiate `purpose` and `abstract` (100% overlap).
Field A Field B Overlap
purpose abstract 100%
purpose keyFacts 49%
abstract keyFacts 49%

Structural size (secondary)

Structural size is secondary. A larger AIPM is not a failure if per-question retrieval stays tiny — check Semantic Redundancy instead.

HTML · 6,325 chars
Schema · 1,559 chars
AIPM · 7,594 chars

Routing helper (secondary)

Illustrative card-first vs HTML-always simulation — prefer per-question measured cost above. Not a ranking.

Card-first retrieval is roughly token-neutral vs HTML-always on this pack.

Orientation pack

7 questions · all layers scored

  • HTML 5/7
  • Schema 1/7
  • AIPM 6/7

Depth pack

3 questions · all layers scored

  • HTML 3/3
  • Schema 1/3
  • AIPM 0/3

Full-context pack metrics (secondary)

These measure the whole file fed to the model this run — not the minimum slice needed per question.

HTML

8/10 matched

21,080 full-pack tokens

Schema

2/10 matched

5,496 full-pack tokens

AIPM

6/10 matched

22,617 full-pack tokens

Full-pack resource table (secondary)

Whole-file context fed this run. Prefer Minimal Retrieval Cost on each question.

Metric HTML Schema AIPM
Coverage (matched) 8/10 2/10 6/10
Context size 6,325 chars 1,559 chars 7,594 chars
Total tokens 21,080 5,496 22,617
Tokens / correct answer 2,635 2,748 3,770
Est. cost / correct answer $0.000406 $0.000438 $0.000574
Matched per 1k tokens 0.380 0.364 0.265
Median latency 987 ms 774 ms 796 ms
Est. cost (USD) $0.00325 $0.00088 $0.00344

Six independent scores

Answer Efficiency is the primary cost lens. Accuracy axes remain for research — no combined total or winner.

Answer Efficiency

Matched answers per 1k tokens (and cost per match). The primary efficiency axis — not raw accuracy.

  • HTML 100
  • Schema 95.8
  • AIPM 69.7

Understanding

Can this layer convey what the page is about — using independent HTML gold?

  • HTML 60
  • Schema 20
  • AIPM 100

Retrieval

Can this layer surface shared facts (location, contact, hours, pricing, FAQ, CTA)?

  • HTML 100
  • Schema 33.3
  • AIPM 0

Evidence

Answer quality vs independent gold (score strength). Partial credit counts; UNKNOWN scores zero unless gold is UNKNOWN.

  • HTML 68.8
  • Schema 25.6
  • AIPM 47.1

Metadata

Language, page kind, and freshness from page signals.

  • HTML 100
  • Schema 0
  • AIPM 50

Compression

Information delivered per token and context size. Higher means more matched answers for less context cost.

  • HTML 100
  • Schema 95.8
  • AIPM 69.7

Coverage map

AIPM matched 6 question(s) (alone on Q3); HTML matched 8. HTML/Schema (or a gap) still needed on Q6, Q8, Q9, Q10. Read this as complementary coverage — AIPM orients agents cheaply; HTML supplies depth when the sidecar cannot. HTML 8/10 matched (≈2,635 tok/match). Schema 2/10 matched (≈2,748 tok/match). AIPM 6/10 matched (≈3,770 tok/match). Figures are descriptive per layer — not a ranking. See per-question chains for sufficient layers, slice size, and measured cost.

AIPM matched

Q1, Q2, Q3, Q4, Q5, Q7

Alone: Q3

AIPM insufficient

Q6, Q8, Q9, Q10

HTML/Schema needed or all layers missed

Unanswered by all

HTML layer

Visible page text after stripping AIPM sidecars and JSON-LD. Measures what prose alone can answer.

Context fed: 6,325 chars

Tokens: 21,080 · median 987 ms

Matched this pack: 8/10

Stronger on

Retrieval (100) · Metadata (100) · Compression (100) · Answer Efficiency (100)

Weaker on

Schema layer

JSON-LD structured data with minimal page chrome. Measures what schema markup can answer.

Context fed: 1,559 chars

Tokens: 5,496 · median 774 ms

Matched this pack: 2/10

Stronger on

Compression (95.8) · Answer Efficiency (95.8)

Weaker on

Understanding (20) · Retrieval (33.3) · Metadata (0) · Evidence (25.6)

AIPM layer

AI Page Manifest (.ai.json) only. Measures what the machine layer can answer without HTML.

Context fed: 7,594 chars

Tokens: 22,617 · median 796 ms

Matched this pack: 6/10

Stronger on

Understanding (100)

Weaker on

Retrieval (0)

When to use which layer

AIPM complements HTML — it does not replace full-page prose.

Scenario Recommended Why
Fast orientation (title, purpose, brand, intent) Compare machine card → HTML fallback on this run Machine-card pack: 22,617 tok · $0.00344.
Deep content / research (prose facts, process detail) HTML (with optional machine orientation) HTML pack: 21,080 tok · $0.00325. Use when depth needs body prose.
Structured entity pulls (org, location, typed fields) Schema.org Schema pack: 5,496 tok · $0.00088. Dense JSON-LD tends to score well here.

Findings

  • Of 10 questions: lowest measured retrieval cost among sufficient layers — HTML 3 · Schema 1 · AIPM 6 · multi-layer 0 · unanswered 0. Descriptive counts only; no overall ranking.
  • Semantic Redundancy 66% — the manifest repeats the same ideas across fields. This inflates structural size without helping per-question retrieval.
  • Manifest design: Merge or differentiate `purpose` and `abstract` (100% overlap).
  • Structural size is secondary. A larger AIPM is not a failure if per-question retrieval stays tiny — check Semantic Redundancy instead.
  • AIPM matched 6 question(s) (alone on Q3); HTML matched 8. HTML/Schema (or a gap) still needed on Q6, Q8, Q9, Q10. Read this as complementary coverage — AIPM orients agents cheaply; HTML supplies depth when the sidecar cannot. HTML 8/10 matched (≈2,635 tok/match). Schema 2/10 matched (≈2,748 tok/match). AIPM 6/10 matched (≈3,770 tok/match). Figures are descriptive per layer — not a ranking. See per-question chains for sufficient layers, slice size, and measured cost.

Question-by-question layer analysis

Which layer(s) could answer; which need more or different context; minimum context fed this run.

Q1 · Understanding · orientation

What is the page title?

Gold: Tij Galvaniz – Kapadokya Galvaniz

Gold source: html.title · html_independent

Sufficient: Schema, AIPM · Needs more / other context: HTML

Estimated retrieval cost · lowest cost AIPM · 3 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.title est no
Schema jsonld.name 5 est yes
AIPM title 3 est yes

HTML

insufficient

UNKNOWN.

score 0 · 6,325 chars context · 2,083 in-tokens

HTML did not answer from 6325 chars of context — additional or different layer context needed.

Schema

sufficient

The page title is "Tij Galvaniz – Kapadokya Galvaniz."

score 89 · 1,559 chars context · 532 in-tokens

Schema answered using 1559 chars of layer context (minimum fed this run).

AIPM

sufficient

The page title is "Tij Galvaniz."

score 67 · 7,594 chars context · 2,244 in-tokens

AIPM answered using 7594 chars of layer context (minimum fed this run).

Q2 · Understanding · orientation

What is the primary topic of this page?

Gold: Tij Galvaniz

Gold source: html.h1 · html_independent

Sufficient: HTML, AIPM · Needs more / other context: Schema

Estimated retrieval cost · lowest cost AIPM · 3 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.h1|title 3 est yes
Schema No typed Schema field for this question. est no
AIPM primaryTopic 3 est yes

HTML

sufficient

The primary topic of this page is Tij Galvaniz and its galvanizing services.

score 88 · 6,325 chars context · 2,086 in-tokens

HTML answered using 6325 chars of layer context (minimum fed this run).

Schema

insufficient

The primary topic of this page is "Kapadokya Galvaniz."

score 33 · 1,559 chars context · 535 in-tokens

Schema did not answer from 1559 chars of context — additional or different layer context needed.

AIPM

sufficient

The primary topic of this page is Tij Galvaniz.

score 88 · 7,594 chars context · 2,247 in-tokens

AIPM answered using 7594 chars of layer context (minimum fed this run).

Q3 · Understanding · orientation

What is the purpose of this page?

Gold: Tijler İçin Galvanizleme Uygulamaları Tij galvaniz, özellikle şu alanlarda yaygın olarak kullanılır ⟶ Endüstriyel Uygulamalar Fabrika ve üretim tesislerinde kullanılan tişler.

Gold source: html.prose · html_independent

Sufficient: AIPM · Needs more / other context: HTML, Schema

Estimated retrieval cost · lowest cost AIPM · 40 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lead_paragraphs est no
Schema jsonld.description est no
AIPM purpose 40 est yes

HTML

insufficient

The purpose of this page is to provide information about Tij Galvaniz services offered by Kapadokya Galvaniz, including galvanizing processes and their advantages.

score 33 · 6,325 chars context · 2,085 in-tokens

HTML did not answer from 6325 chars of context — additional or different layer context needed.

Schema

insufficient

UNKNOWN.

score 0 · 1,559 chars context · 534 in-tokens

Schema did not answer from 1559 chars of context — additional or different layer context needed.

AIPM

sufficient

The purpose of this page is to provide durable and reliable solutions for Tij galvaniz in industrial applications, the automotive sector, construction projects, and more.

score 70 · 7,594 chars context · 2,246 in-tokens

AIPM answered using 7594 chars of layer context (minimum fed this run).

Q4 · Understanding · orientation

Summarize the page in one sentence.

Gold: Tijler İçin Galvanizleme Uygulamaları Tij galvaniz, özellikle şu alanlarda yaygın olarak kullanılır ⟶ Endüstriyel Uygulamalar Fabrika ve üretim tesislerinde kullanılan tişler.

Gold source: html.meta_description|prose · html_independent

Sufficient: HTML, AIPM · Needs more / other context: Schema

Estimated retrieval cost · lowest cost AIPM · 40 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lead_paragraphs 176 est yes
Schema jsonld.description est no
AIPM abstract 40 est yes

HTML

sufficient

Kapadokya Galvaniz offers high-quality galvanizing services for various industrial applications, enhancing the durability and corrosion resistance of components like tijler.

score 70 · 6,325 chars context · 2,086 in-tokens

HTML answered using 6325 chars of layer context (minimum fed this run).

Schema

insufficient

Kapadokya Galvaniz, founded in 2023 by Mahmut Şimşek and Mustafa Gümüş in Kahramankazan, Ankara, specializes in hot-dip centrifugal galvanizing and is associated with Desum Enerji, which has been operational since 2014.

score 33 · 1,559 chars context · 535 in-tokens

Schema did not answer from 1559 chars of context — additional or different layer context needed.

AIPM

sufficient

The page discusses Kapadokya Galvaniz's Tij Galvaniz solutions, highlighting their durability and reliability for industrial applications, automotive sectors, and construction projects.

score 70 · 7,594 chars context · 2,247 in-tokens

AIPM answered using 7594 chars of layer context (minimum fed this run).

Q5 · Understanding · orientation

Who is the publisher or brand?

Gold: Kapadokya Galvaniz

Gold source: html.jsonld|title · html_independent

Sufficient: HTML, AIPM · Needs more / other context: Schema

Estimated retrieval cost · lowest cost AIPM · 5 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.title_brand 9 est yes
Schema jsonld.name est no
AIPM publisher 5 est yes

HTML

sufficient

Kapadokya Galvaniz.

score 100 · 6,325 chars context · 2,084 in-tokens

HTML answered using 6325 chars of layer context (minimum fed this run).

Schema

insufficient

UNKNOWN.

score 0 · 1,559 chars context · 533 in-tokens

Schema did not answer from 1559 chars of context — additional or different layer context needed.

AIPM

sufficient

The publisher is Kapadokya Galvaniz.

score 88 · 7,594 chars context · 2,245 in-tokens

AIPM answered using 7594 chars of layer context (minimum fed this run).

Q6 · Metadata · orientation

What language is this page in?

Gold: tr

Gold source: html.lang|script · html_independent

Sufficient: HTML · Needs more / other context: Schema, AIPM

Estimated retrieval cost · lowest cost HTML · 2 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.lang 2 est yes
Schema jsonld.block est no
AIPM inLanguage est no

Root cause · Layer miss

Improve orientation fields; do not paste full HTML into the sidecar.

HTML

sufficient

The page is in Turkish.

score 88 · 6,325 chars context · 2,084 in-tokens

HTML answered using 6325 chars of layer context (minimum fed this run).

Schema

insufficient

UNKNOWN.

score 0 · 1,559 chars context · 533 in-tokens

Schema did not answer from 1559 chars of context — additional or different layer context needed.

AIPM

insufficient

The page is in English.

score 0 · 7,594 chars context · 2,245 in-tokens

AIPM did not answer from 7594 chars of context — additional or different layer context needed.

Q7 · Metadata · orientation

What is the content intent?

Gold: commercial

Gold source: html.heuristic · html_independent

Sufficient: HTML, AIPM · Needs more / other context: Schema

Estimated retrieval cost · lowest cost AIPM · 3 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.prose_signals 100 est yes
Schema No typed Schema field for this question. est no
AIPM contentIntent 3 est yes

HTML

sufficient

The content intent is to provide information about Tij Galvaniz services and the benefits of galvanizing products for industrial applications.

score 88 · 6,325 chars context · 2,083 in-tokens

HTML answered using 6325 chars of layer context (minimum fed this run).

Schema

insufficient

UNKNOWN.

score 0 · 1,559 chars context · 532 in-tokens

Schema did not answer from 1559 chars of context — additional or different layer context needed.

AIPM

sufficient

The content intent is commercial.

score 88 · 7,594 chars context · 2,244 in-tokens

AIPM answered using 7594 chars of layer context (minimum fed this run).

Q8 · Retrieval · depth

Which location or city is mentioned for this business?

Gold: Kahramankazan

Gold source: html.address|prose · html_independent

Sufficient: HTML, Schema · Needs more / other context: AIPM

Estimated retrieval cost · lowest cost SCHEMA · 4 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.address|prose 4 est yes
Schema jsonld.addressLocality 4 est yes
AIPM Field missing in AIPM — cannot answer from this layer alone. est no

Root cause · Depth gap (expected)

Prose/detail questions belong to HTML — escalate after AIPM orientation.

HTML

sufficient

Kahramankazan / ANKARA.

score 88 · 6,325 chars context · 2,087 in-tokens

HTML answered using 6325 chars of layer context (minimum fed this run).

Schema

sufficient

Kahramankazan.

score 100 · 1,559 chars context · 536 in-tokens

Schema answered using 1559 chars of layer context (minimum fed this run).

AIPM

insufficient

UNKNOWN.

score 0 · 7,594 chars context · 2,248 in-tokens

AIPM did not answer from 7594 chars of context — additional or different layer context needed.

Q9 · Retrieval · depth

According to the page, what does it say about: "Asitleme: Temizlenen demirler, asit banyosunda k…"?

Gold: Asitleme: Temizlenen demirler, asit banyosunda korozyon ürünlerinden arındırılır.

Gold source: html.prose · html_independent

Sufficient: HTML · Needs more / other context: Schema, AIPM

Estimated retrieval cost · lowest cost HTML · 21 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.prose_span 21 est yes
Schema Schema present but no minimal typed field for this question. est no
AIPM Prose fact not present as a discrete AIPM field — escalate to HTML. est no

Root cause · Depth gap (expected)

Prose/detail questions belong to HTML — escalate after AIPM orientation.

HTML

sufficient

Asitleme, temizlenen demirlerin asit banyosunda korozyon ürünlerinden arındırılması işlemidir.

score 71 · 6,325 chars context · 2,108 in-tokens

HTML answered using 6325 chars of layer context (minimum fed this run).

Schema

insufficient

UNKNOWN.

score 0 · 1,559 chars context · 557 in-tokens

Schema did not answer from 1559 chars of context — additional or different layer context needed.

AIPM

insufficient

UNKNOWN

score 0 · 7,594 chars context · 2,269 in-tokens

AIPM did not answer from 7594 chars of context — additional or different layer context needed.

Q10 · Retrieval · depth

According to the page, what does it say about: "Çinko Kaplama: Demirler, erimiş çinko banyosuna …"?

Gold: Çinko Kaplama: Demirler, erimiş çinko banyosuna daldırılarak kaplanır.

Gold source: html.prose · html_independent

Sufficient: HTML · Needs more / other context: Schema, AIPM

Estimated retrieval cost · lowest cost HTML · 18 input tokens

Layer Planner fields Input tokens Rounds Confidence OK?
HTML html.prose_span 18 est yes
Schema Schema present but no minimal typed field for this question. est no
AIPM Prose fact not present as a discrete AIPM field — escalate to HTML. est no

Root cause · Depth gap (expected)

Prose/detail questions belong to HTML — escalate after AIPM orientation.

HTML

sufficient

Çinko kaplama işlemi, demirlerin erimiş çinko banyosuna daldırılarak kaplanmasını içerir.

score 63 · 6,325 chars context · 2,108 in-tokens

HTML answered using 6325 chars of layer context (minimum fed this run).

Schema

insufficient

UNKNOWN.

score 0 · 1,559 chars context · 557 in-tokens

Schema did not answer from 1559 chars of context — additional or different layer context needed.

AIPM

insufficient

UNKNOWN

score 0 · 7,594 chars context · 2,269 in-tokens

AIPM did not answer from 7594 chars of context — additional or different layer context needed.

Methodology

  • AI Context Benchmark: Planner → Slice → LLM with measured API input tokens.
  • Reports describe sufficient layers, slice size, cost, and evidence — they do not declare a winning format.
  • Layers under test today: HTML, Schema.org, AIPM (extensible to RSS, Markdown, PDF, …).
  • Engine aipm_benchmark_v3 · Wed, Jul 29, 2026 10:31 AM · aipm_benchmark_score_v10

Share

Per-question context chain across layers.