Home · Journal · Digital marketing
Digital marketing24 min read11/09/2026

Your site answers Google. It often refuses ChatGPT

We took every OpenStreetMap unit in Cluj county — restaurants, hotels, cabins, cafés, dental practices, medical practices and law firms — that had filled in a website on the map, and asked each homepage four times from the same address on the same day, changing only the name we introduced ourselves with. 231 answered, and 228 made it through the full test. A fake Googlebot was refused by 15 of them. The crawler that trains ChatGPT, by 97. The one that feeds ChatGPT search, by 71 out of 225.

Black and white modern facade, a grid of identical closed windows with a single one open, reflecting the sky
One open window in a wall of identical ones. That is roughly the difference between a site that can be read and the rest.
In short

A visit from a search engine used to be a visit from one kind of program. It is now several: crawlers that index, crawlers that train models, and fetchers that run because a person just asked a question. Each arrives under a different name, and a lot of websites treat those names very differently — usually without anybody having decided to.

This is not a story about doing more marketing. It is a list of things that either work on your website or do not, measured rather than assumed, on 231 real Romanian business sites. Most of the expensive advice in this field sits on top of a foundation that, in a third of cases, is not there.

Below: the measurement and the rule that defined it, the four gates a page has to pass before any content work is worth paying for, the full process from zero to steady growth — clean methods only — and the one thing nobody honest can guarantee you.

  • You will know whether the systems that now answer instead of the search engine can read you at all.
  • You will know what takes hours to fix and what takes months, and the order between them.
  • You will know which numbers to watch, in which order, and which ones are theatre.
  • You will know what this measurement does not prove, before somebody else tells you.
What we found, in one place. Cittago measurement, Cluj county, 11 September 2026
What we measuredResult
Units mapped in OpenStreetMap1,782
Unique domains after deduplication314
Sites that answered231
Refuse GPTBot (OpenAI training)97 of 228 — 42.5%
Refuse OAI-SearchBot (ChatGPT search)71 of 225 — 31.6%
Refuse ClaudeBot79 of 228 — 34.6%
Refuse a fake Googlebot15 of 228 — 6.6%
Wrote the block in robots.txt11 of 97
Have local business structured data17 of 231 — 7.4%
Have sitemap.xml55.0%
Have llms.txt14.7%
Have at least one h166.2%
Median time to full HTML475 ms

What we measured, and the rule that defined it

Most writing about search optimisation is opinion with a confident tone. We wanted numbers, so we defined a set by a rule, measured it, and are publishing the result whichever way it came out. Here is the rule, so anyone can repeat it.

The set: every OpenStreetMap unit inside Cluj county, Romania (administrative relation 91733), across seven categories — restaurants; cafés, bars and fast food; hotels and motels; cabins, guesthouses and hostels; dental practices; medical practices and clinics; law firms — that had a website address filled in on the map. Snapshot taken on 11 September 2026.

That produced 1,782 mapped units, of which 347 had a website of their own and 15 had only a social profile in place of one. After deduplicating by domain, 314 unique domains remained. Of those, 231 answered with a page; 80 did not answer even on a second attempt, and 3 recovered on retest.

What this set does not prove, written before you read the numbers: a unit without a website in OpenStreetMap does not mean the business has no website. It means nobody filled it in on the map. So the percentages below do not read as “this share of Cluj restaurants have a website”. They read as “this is what the Cluj websites we can actually check look like”. It is a convenience sample, not a representative one — and its only real guarantee is that we did not choose who went into it.

For each domain we measured: whether it answers and how fast, whether it lands on HTTPS, what its robots.txt says, whether a sitemap exists at the standard address, whether llms.txt exists, what structured data the page carries, whether it has a title, a description, an h1 and a mobile viewport tag, what it is built on — and, the part that turned out to matter most, how it answers identical requests made under four different names.

No domain name is published. There is no ranking and no competitor audit here. Aggregate figures, and nothing else.

02 · What the page is missing

Almost all have HTTPS. One in thirteen says who it is

The certificate and the mobile viewport sorted themselves out, through themes and hosting. What has to be written by hand — who you are, where you are, when you are open — is missing almost everywhere.

7.4%have local business schema
33.8%have no h1 heading at all
0%25%50%75%100%2.2%FAQPage schema7.4%local business schema14.7%llms.txt55.0%sitemap.xml64.9%meta description66.2%at least one h176.6%valid robots.txt93.5%mobile viewport94.8%HTTPS

Source: Cittago’s own measurement, 11 September 2026. Set: OpenStreetMap units in Cluj county (relation 91733), seven categories, with a website filled in on the map. 314 unique domains, 231 responded. cittago.com

Almost all have HTTPS. One in thirteen says who it is
value
FAQPage schema2.2%
local business schema7.4%
llms.txt14.7%
sitemap.xml55.0%
meta description64.9%
at least one h166.2%
valid robots.txt76.6%
mobile viewport93.5%
HTTPS94.8%
Local business structured data requires two mandatory fields: name and address. 17 sites out of 231 have them.

Read top to bottom, the chart says something we see well beyond Cluj: what solved itself, has solved itself. HTTPS comes from the hosting. The mobile viewport comes from the theme. Almost everyone has them, without anyone having decided anything.

What has to be written by hand is missing. Local business structured data — the block that tells a machine, in text it does not have to guess at, what you are called, where you are, when you open and what kind of business you run — exists on 17 sites out of 231. Google requires exactly two mandatory fields for it, name and address, and asks for the most specific subtype available: not a generic “LocalBusiness”, but “Restaurant”, “Dentist”, “LegalService”. Seven in a hundred have it.

An h1 heading — the simplest thing in the whole document, the sentence that says what this page is — is missing from 78 sites out of 231. A third. A meta description is missing from 81. These are not specialist subtleties; they are the fields any modern theme fills in the moment someone writes the text for them.

Pale abstract render of an artificial intelligence system: small coloured geometric shapes arranged on a white surface
A model does not “see” the internet. It receives text, from programs that call at your house and leave with whatever you gave them.

Gate one: can anything read you?

Before any keyword, any copy and any link, there is a mechanical question: when a program arrives to read your page, does it get the page or a closed door?

We asked each homepage four times, from the same server, on the same day, within the same hour. The requests were identical. The only thing that changed between them was the name we introduced ourselves with — the user-agent string.

01 · Who gets through the gate

Same request, different name: Googlebot passes, GPTBot does not

We asked each site’s homepage four times from the same IP address, on the same day. The only thing that changed between requests was the user-agent string. A fake Googlebot is refused by 15 sites. A GPTBot is refused by 97.

42.5%refuse GPTBot
83of those let Googlebot through
11wrote it in robots.txt
02040606.6%Googlebot (fake)8.3%curl/8.5.034.6%ClaudeBot42.5%GPTBot

Source: Cittago’s own measurement, 11 September 2026. Set: OpenStreetMap units in Cluj county (relation 91733), seven categories, with a website filled in on the map. 314 unique domains, 231 responded. cittago.com

Same request, different name: Googlebot passes, GPTBot does not
value
Googlebot (fake)6.6%
curl/8.5.08.3%
ClaudeBot34.6%
GPTBot42.5%
All four requests left the same server within the same hour. The difference is not in the network, it is in a list.

A plain `curl` is refused by 19 sites out of 228. A fake Googlebot, by 15. A ClaudeBot, by 79. A GPTBot, by 97.

The obvious explanation is that these sites run generic anti-robot defences and reject anything that is not a browser. That is exactly why we added the controls. If it were true, the same sites would also refuse `curl` and the fake Googlebot. Of the 97 that block GPTBot, only 13 do. Another 83 let the fake Googlebot straight through and stop only GPTBot. That is not a generic defence. That is a list.

And here is the uncomfortable part. Of those 97, only 11 wrote anything about AI crawlers in `robots.txt`. The other 86 block without the owner having written a line. What GPTBot receives instead of the page: 58 times a 403, 19 times a 429, 5 times a 410, 3 times a 406, and 11 times the connection simply fails. Only 11 of the 97 sit behind Cloudflare, so the blocking cannot be pinned on it. It comes from the security modules that ship by default with ordinary hosting.

The asymmetry points to inheritance rather than decision: 28 sites block GPTBot but let ClaudeBot through, and 10 do the exact opposite. A considered policy would produce similar lists. These are different lists because they come from different defaults, written at different times.

But GPTBot is not the crawler that puts you inside ChatGPT answers. GPTBot collects pages to train models. Search inside ChatGPT runs under a different name, `OAI-SearchBot`, with a different list and a different decision behind it. A site can refuse one and allow the other. We had not measured that, so we ran a second round the same day, on the same sites.

Second round, 11 September 2026: the 225 sites that also answered a browser in this run
CrawlerWhat it doesRefused byShare
GPTBottraining OpenAI models94 of 22541.8%
OAI-SearchBotsearch inside ChatGPT71 of 22531.6%
PerplexityBotPerplexity search78 of 22534.7%

The overlap says more than the percentages. 68 sites refuse both OpenAI crawlers, training and search alike. 26 refuse training only and let search through. 3 do the exact opposite, which helps nobody.

Those 26 are the only ones in the whole set that look like a position somebody took: do not train on my text, but do quote me. It is also what we recommend when anyone asks. The other 68 separate nothing.

The second round also checked the first. Across the 224 sites common to both runs, GPTBot was refused 94 times in each, with the same verdict on 93.8% of them. The differences are sites that went down or came back between runs, not a number shifting under our feet.

03 · By line of business

Hotels block the most, dental practices the least

The share of sites that answer a browser but not a GPTBot, across the seven categories. Small samples — law firms have six sites, dentistry twelve — read as an indication, not a measurement.

0%20%40%60%25.0%Dental practices (12)32.3%Cafés, bars (31)33.3%Law firms (6)42.5%Restaurants (73)44.8%Medical practices (29)45.2%Cabins, guesthouses (31)52.2%Hotels (46)

Source: Cittago’s own measurement, 11 September 2026. Set: OpenStreetMap units in Cluj county (relation 91733), seven categories, with a website filled in on the map. 314 unique domains, 231 responded. cittago.com

Hotels block the most, dental practices the least
value
Dental practices (12)25.0%
Cafés, bars (31)32.3%
Law firms (6)33.3%
Restaurants (73)42.5%
Medical practices (29)44.8%
Cabins, guesthouses (31)45.2%
Hotels (46)52.2%
The number of sites in each category is in the table. Below ten, the percentage says more about the sample than about the market.

What robots.txt is, and what it is not

It is a text file at the root of the domain that asks programs passing by not to read certain parts. The important word is “asks”. It is a convention, not an access control. Cloudflare says it plainly about its own content signals: “content signals express preferences; they are not technical countermeasures against scraping”.

Who honours it, according to each operator’s own documentation, on 11 September 2026:

What each operator states about its own crawler, in its own documentation
BotOperatorPurposeHonours robots.txt?
GPTBotOpenAImodel trainingyes
OAI-SearchBotOpenAIChatGPT searchyes
ChatGPT-UserOpenAIuser-initiated fetches“may not apply”
ClaudeBotAnthropiccollection for modelsyes
Claude-UserAnthropicuser questionsyes
PerplexityBotPerplexitysurfacing in resultsyes
Perplexity-UserPerplexityuser-initiated fetches“generally ignores”
GooglebotGoogleSearch indexingyes
Google-ExtendedGooglepolicy token onlydoes not crawl at all
meta-externalagentMetatraining and indexingyes
meta-externalfetcherMetaagentic tasks“may bypass”

The fault line does not run between companies, as you might expect. It runs between automated crawlers, which honour the rules, and fetches triggered by a person, where three operators out of three write in their own documentation that the rules may not apply. Perplexity is the most explicit: “Since a user requested the fetch, this fetcher generally ignores robots.txt rules.”

Two traps we see constantly. First: Google-Extended and Applebot-Extended are not robots. They are names you can use in robots.txt to say “do not use me for training”. They never appear in your access logs, because they never request anything. Blocking Google-Extended does not reduce crawling and does not affect indexing in Search. Second: Applebot follows instructions given to Googlebot if robots.txt does not mention it. A `Disallow` aimed at Googlebot quietly closes the door on Siri and Spotlight too.

The security modules that block without asking you

This is where the 86 come from. Ordinary hosting ships with a protection module — a web application firewall — running a rule set over every request. The most widespread set, the OWASP Core Rule Set, has four aggressiveness levels and works by accumulating a score: no single rule blocks on its own, but once the accumulated score reaches the recommended inbound threshold of 5, the transaction is refused. Its own documentation admits that “when working in strict blocking mode, false positives can cause legitimate user transactions to be blocked”.

Add rate limiting, JavaScript challenges, and vendor-maintained user-agent lists, and you get a site that blocks crawlers without anyone in the company having touched a setting. The consequence is not limited to AI systems. Google documents that it actively reduces its crawl rate when it receives 5xx errors or rate-limiting signals such as HTTP 429: “the limit goes down and Google crawls less”. A badly tuned rate limiter does not delay one page; it shrinks the crawl budget of the entire site.

In Search Console this shows up under a dedicated state — Blocked due to access forbidden (403) — which Google defines as the case where “the user agent provided credentials, but was not granted access”.

The decision is yours, not your host’s

We are not saying you must let AI crawlers in. It is a legitimate business decision in both directions. A publisher living on advertising has good reasons to refuse training. A restaurant that wants to be recommended when someone asks an assistant where to eat has good reasons to accept.

The problem is different: those 86 decided nothing. They inherited a setting. Cloudflare drew the distinction that matters in its content signals policy, published on 24 September 2025: `ai-train` covers training models, and `ai-input` covers using content in real time as the source of an answer. They are two separate permissions, and you may well want one without the other — to be cited without being trained on.

From 15 September 2026 Cloudflare applies new defaults for newly onboarded domains: bots classified as Training and Agent are blocked by default on pages that display ads, while Search stays allowed. That condition about ads is essential and disappears from almost every summary of the change.

sitemap.xml: what it does and what it does not

A sitemap is a list of addresses you hand the engine so it does not depend on links alone. The official specification caps a file at 50,000 URLs and 50 MB uncompressed. Google confirms it helps large or complex sites, but does not guarantee that everything in it will be crawled and indexed, and says a small site — around 500 pages — that is well linked internally may not need one at all.

The part most often got wrong is `lastmod`. In the specification it is the date the page last changed. Many systems write the date the sitemap was generated, for every address at once. Google is explicit: it uses the value “if it’s consistently and verifiably accurate”, and ignores `priority` and `changefreq` entirely. A false `lastmod` does not merely fail to help — it makes Google distrust `lastmod` across the whole site. Changing the copyright year in the footer does not count as a change; changing the main content, the structured data or the links does.

llms.txt: what it is, and what we do not know about it

It was proposed by Jeremy Howard on 3 September 2024: a Markdown file at `/llms.txt`, written to be read by language models. The only mandatory section is an H1 with the project name. The specification has been revised; the official page shows 10 August 2026 as its modification date. It remains a proposal open to community input, not a standard ratified by any body.

14.7% of our set has it — 34 sites out of 231 — which is more than we expected and almost certainly comes from themes and plugins generating it automatically, not from decisions.

What we do not know, and it is right to say so: whether anyone reads it. OpenAI’s crawler documentation does not mention it at all and points exclusively to robots.txt. Anthropic publishes its own llms.txt files, but publishing does not prove consumption. Google has the closest thing to an official statement, in a guide last updated 10 July 2026: “You don’t need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search” — without naming llms.txt. So: cheap to add, no evidence of effect. We have one, for the same reason we leave the windows unlocked as well as the door.

Search Console: registration, and the errors that matter

If you do not have the property verified in Google Search Console, you cannot know any of what follows, and any report about your site is a story. It is free and takes minutes.

The indexing states you will see there, with what each means in Google’s documentation:

  • Discovered — currently not indexed: Google found the address but has not read it yet. The documented typical reason is that a crawl “was expected to overload the site”. So it is a server capacity problem, not a copy quality problem — which ties this state directly to the next section.
  • Crawled — currently not indexed: it read the page and did not index it. Google says explicitly “no need to resubmit this URL for crawling”, and does not give you the reason.
  • URL is unknown to Google: literally, Google has never seen the address. That is a discovery failure — sitemap, internal links, blocked access — not an indexing failure. Different remedy.
  • Indexed, though blocked by robots.txt: the logical trap. robots.txt blocks reading, not indexing. A page blocked there cannot be read, so a `noindex` written inside it can never be seen.
  • Page indexed without content: the page is in the index but Google could not read the content. This is the state you get when a crawler is served a JavaScript challenge or an interstitial instead of text.
  • Soft 404: the page says “not found” but answers with status 200.

These six are fixed in six different ways, which is why “I have errors in Search Console” is not a diagnosis. The SEO and AI search side of our work starts exactly here, by reading the states — because there is no point writing new copy for a page Google has never found.

Metal racks filled with sheets of coloured paper, photographed head-on in low light
A set defined by a rule looks like this: everything goes into the same drawers, then you count. You do not pick who goes in.

SEO and GEO: what each word means, and since when

The two words now travel together and get confused constantly, including inside proposals. They are worth separating, because they ask for different things.

SEO — search engine optimisation — is the work of making a site easy for a search engine to find, read and rank, so it appears on the results page when someone searches for something you offer. The term is nearly thirty years old and its authorship is disputed: Bruce Clay Inc. has been positioning sites since January 1996 and claims on its own site the credit for first using the phrase, but we could not open a single dated artefact from 1996 or 1997 that actually contains it. Other names in the same dispute are John Audette, Bob Heyman and Danny Sullivan. Google itself was founded in August 1998, after Andy Bechtolsheim wrote a cheque for 100,000 dollars.

GEO — generative engine optimisation — is the work of making a site easy for a system that answers, rather than listing links, to read, understand and cite. Its origin is far more precise than SEO’s: the academic paper “GEO: Generative Engine Optimization” was submitted to arXiv on 16 November 2023 by six authors from IIT Delhi and Princeton, and was accepted at KDD 2024. The authors built a benchmark of 10,000 queries and report visibility gains “of up to 40%” — with their own explicit caveat that this is a maximum rather than an average, and that it is measured on two metrics they defined, not on real traffic.

The two disciplines, on the things that matter in practice
SEOGEO
Since whendisputed, around 1996–199716 November 2023, academic paper
Where you appearin the list of resultsinside a generated answer
What gets you thererelevance, links, page experiencebeing accessible, quotable and checkable
What you measureimpressions, clicks, position, in Search Consolementions and citations, counted by hand
Time to first signalmonthsmonths, and with no official dashboard

Why GEO became a concern only recently: because generated answers moved from experiment to default. Google showed its first generative capabilities in search on 10 May 2023, as an opt-in experiment in Search Labs. On 14 May 2024, AI Overviews began rolling out to everyone in the United States. On 28 October 2024 it reached more than 100 countries, and on 20 May 2025 more than 200 countries and over 40 languages. In parallel, OpenAI launched search in ChatGPT on 31 October 2024, opened it to all signed-in users on 16 December 2024, and removed the account barrier entirely on 5 February 2025.

What happened to clicks afterwards is contested, and deserves to be presented as contested rather than settled. Pew Research Center published a study on 22 July 2025 covering 68,879 searches by 900 US adults in March 2025: users clicked a traditional result in 8% of visits where an AI summary appeared, against 15% where it did not. They clicked a source inside the summary in 1% of visits. Ahrefs, which sells SEO tools and therefore has a commercial interest in the answer, measured a 34.5% lower CTR on position one where an AI Overview appears in April 2025, and repeated the study in February 2026 at −58%.

On the other side, Liz Reid, who runs Google Search, wrote on 6 August 2025 that “total organic click volume from Google Search to websites has been relatively stable year-over-year”. Google publishes no figure, no denominator and no methodology in support, while accusing third-party studies of flawed methodologies. The two camps do not fully contradict each other: Google is talking in aggregate, across the whole web, and admits some sites lose while others gain. That says nothing about yours.

The practical conclusion, the only one that matters for a business: it is no longer enough for your page to be found by a person looking at a list. It has to be readable by a machine composing an answer.

Brick wall with an intercom mounted beside a half-open metal door
The intercom either answers or it does not. The difference here is that you never hear the buzzer, and nobody tells you somebody called.

Gate two: how long a visitor waits

If the page can be read, the next question is how long it takes to appear. There is an official measurement here, with published thresholds, and a mountain of mythology on top of it.

What is actually measured in 2026

Core Web Vitals are three metrics, not more. On 11 September 2026, Google’s documentation lists exactly that, with unchanged thresholds:

Core Web Vitals and their thresholds, verified in Google’s documentation on 11 September 2026
MetricWhat it measuresGoodPoor
LCPwhen the largest visible element appearsunder 2.5 sover 4.0 s
INPhow long a reaction to an interaction takesunder 200 msover 500 ms
CLShow much content jumps while loadingunder 0.1over 0.25

INP officially replaced FID on 12 March 2024. All three are assessed at the 75th percentile of page loads, separately on mobile and desktop — not at the average. Which means a “good” site still leaves a quarter of its visitors with a worse experience than the number it reports.

A warning, because it is in season: articles are circulating in volume about new Core Web Vitals metrics introduced in 2026, about a “Visual Stability Index”, or about the LCP threshold dropping to 2.0 seconds. We searched Google’s documentation and the Chrome UX Report release notes. There is no primary source for any of it. What actually changed across 2025 and 2026 are diagnostic dimensions, not the metrics and not the thresholds.

Why PageSpeed gives you a different score every time

PageSpeed Insights shows two things that get confused permanently. Field data comes from the Chrome UX Report, over a 28-day collection window, at the 75th percentile — those are your real visitors. Lab data is a test run on the spot, in a controlled environment.

The lab test does not run on a phone. It uses simulated Slow 4G throttling — 1.6 Mbps down, 750 Kbps up, 150 ms latency — plus a constant 4× CPU multiplier. Lighthouse’s own documentation acknowledges an “inherent inaccuracy” in the simulation. The score shown is a weighted average of five metrics, dominated by Total Blocking Time at 30%, LCP at 25% and CLS at 25%. Notice what is missing: INP is not in the lab score, so the PageSpeed number is not the same thing as the Core Web Vitals assessment used in ranking.

Google lists the sources of variability explicitly — local network, client hardware, resource contention — without publishing a numeric range. Which is why our rule is two runs and a control. A single run is not a measurement.

And the 28-day window explains why a fix made today does not fully show up in field data for about four weeks. Anyone promising an improvement “in the real data” within two days does not know how the tool works.

How much speed is worth in money

Three kinds of evidence get cited interchangeably here, and they are not equivalent.

Controlled tests. The strongest and the rarest. Vodafone ran an A/B test on a landing page where the only difference was Web Vitals optimisation: a 31% LCP improvement produced 8% more sales, plus 15% on lead-to-visit rate. Rakuten 24 ran a month at a 50/50 split between an optimised page and the original: +53.37% revenue per visitor, +33.13% conversions. Both are one page, one brand.

Large correlational studies. “Milliseconds Make Millions”, commissioned by Google and run with Deloitte Digital across 37 brands and 30 million mobile sessions over 4 weeks, found on retail that a 0.1 s improvement correlates with +8.4% conversions and +9.2% average order value. The report itself warns that results “may not fully reflect the Internet as a whole”, that it is mobile-only, and that on desktop the team found contradictory parameters. On lead generation — which is exactly the shape of a clinic or a law firm — the result was mixed: progression through the form rose 21.6%, but mobile conversion rate fell by about 2%. Anyone quoting this study without that paragraph is selling, not informing.

Old numbers circulating as eternal truths. “53% of mobile visitors abandon a page that takes over 3 seconds” comes from aggregated Google Analytics data across 3,700 mobile sites, from March 2016. It is over ten years old and the sample was self-selected. We mention it only to say where it comes from.

04 · How long a visitor waits

Half answer in under half a second. One in ten needs more than two

Time to complete HTML, measured from a server in Europe, without a browser and without images. It is the floor under what a real visitor feels on a phone — not their score, the floor beneath it.

29.4%over one second
8.7%over 2.5 seconds
01,0002,0003,000475 msmedian1,105 msthree quarters under2,345 msnine in ten under

Source: Cittago’s own measurement, 11 September 2026. Set: OpenStreetMap units in Cluj county (relation 91733), seven categories, with a website filled in on the map. 314 unique domains, 231 responded. cittago.com

Half answer in under half a second. One in ten needs more than two
value
median475 ms
three quarters under1,105 ms
nine in ten under2,345 ms
This is not Core Web Vitals and does not compare with PageSpeed. It is the server alone, without rendering.

Our own measurement across the 231 Cluj sites is deliberately narrow: time to complete HTML, from a server in Europe, with no browser and no images. It is not Core Web Vitals and does not compare with PageSpeed. It is the floor — what the server does before the browser starts working. The median is 475 ms, three quarters sit under 1,105 ms, and one in ten needs more than 2,345 ms. 68 sites take over a second just to deliver the text.

What Cloudflare fixes, and what it does not

First trap: the Cloudflare CDN does not cache HTML by default. The documentation is direct: “The Cloudflare CDN does not cache HTML or JSON by default.” It caches a fixed list of static extensions. So “we put Cloudflare on it” does not mean your pages are served from cache.

For WordPress, the product that does that is Automatic Platform Optimization: it serves HTML from the edge network, with a 30-day default TTL, bypassing the cache for signed-in users. It requires the official Cloudflare plugin and costs 5 dollars a month on the free plan, included on paid plans.

Second trap, and it is a serious one: Cloudflare refuses by default to cache a response carrying a `Set-Cookie` header. Even with “Cache Everything” on, it keeps the cookie and does not store the page — “a cache MISS will be returned every time”. That is the mechanism keeping cart and checkout uncached. But if someone sets an explicit Edge Cache TTL through rules, Cloudflare strips `Set-Cookie` and caches the page — which is precisely how a shop ends up serving one customer’s basket to another.

What changed and is still being recommended wrongly: Auto Minify was retired on 5 August 2024, and the Brotli toggle was deprecated on 15 August 2024 — compression stayed on, only the switch disappeared. Mirage reached end of life in January 2026. Guides telling you to tick Auto Minify are written for a dashboard that no longer exists.

Argo Smart Routing caches nothing; it picks a less congested network path, and Cloudflare claims 30% better performance on average — a marketing figure with no published methodology. Polish compresses images but does not resize them, so it does not solve a 3,000-pixel image being served to a phone.

Three-dimensional illustration of a hand holding a phone with a lit screen, on a black background
The visitor does not know and does not care whether your server is shared. They only know nothing has happened yet.

Hosting, for a WooCommerce shop

WooCommerce has the largest share of detected e-commerce sites — 44.4% on mobile — and at the same time the worst pass rate: 35% of WooCommerce shops pass all three Core Web Vitals on mobile, against 76% for Shopify and 66% for Wix eCommerce. Broken down, the problem is clear: LCP good at 39%, while INP is good at 88% and CLS at 85%. It is not interactivity. It is how long the page takes to appear.

WordPress overall sits at 45% of sites with good Core Web Vitals on mobile, with a median mobile Lighthouse performance score of 41 out of 100.

Why the hardware matters here, with vendor figures:

  • Dedicated versus shared CPU. On a shared processor, the cycles you get depend on your neighbours: a busy neighbour leaves you “fractions of hyper-threads instead of dedicated access to the underlying physical processors”. The documentation explicitly recommends dedicated CPU where “variable performance is intolerable”.
  • NVMe versus SATA versus spinning disk. At the same cloud vendor, a network SSD gives 30 read IOPS per gibibyte against 0.75 for a spinning disk — a 40-to-1 ratio. A Local SSD on NVMe attached physically to the host reaches 680,000 read IOPS, because it does not cross the network. At the same SSD manufacturer, an NVMe PCIe 4.0 model does 200,000 random write IOPS and 6,800 MB/s sequential read, against 30,000 IOPS and 550 MB/s for the SATA equivalent.
  • RAM. WooCommerce asks, as of September 2026, for PHP 8.3 or newer, MySQL 8.0 or MariaDB 10.6, and a WordPress memory limit of at least 256 MB. That figure is per PHP process. Multiplied by the number of concurrent processes it gives the machine’s real requirement — the link the documentation does not make, but which explains why a shop falls over precisely at its traffic peak.

Redis and LiteSpeed: what they actually do

In WordPress, the object cache is non-persistent by default: data lives in memory only for the duration of a single request and is lost on the next page load. The documentation states the consequence directly: “Without a persistent object cache, your web server must read those options from the database to handle every page view.”

Redis, installed as a persistent backend, keeps results between requests. Note what it does not do: it is an object and query cache, not an HTML page cache. And there is a detail that bites shops: WordPress transients “should also never be assumed to be in the database”. WooCommerce stores computed results in transients — taxes, variations, reports. A full Redis evicting keys produces silent recalculation, not visible errors.

LiteSpeed is a different animal. It is not a PHP caching plugin, it is the caching engine built into the web server — “LiteSpeed’s more efficient and highly customizable answer to Apache mod_cache and Varnish”. The practical difference: a cached page is served without starting PHP. The vendor claims halved server load and 3× better TTFB against Apache; those are marketing figures with no published methodology.

The mechanism that matters for a shop: LiteSpeed separates public cache from private cache, and the private cache key includes session cookies. If the application starts a PHP session on every request, every visitor gets their own key and the public cache hit rate drops to zero, even where the page would be identical for everyone. ESI — Edge Side Includes — solves that by breaking the page into fragments: the public part is cached, while “hello, Ioana” and the basket regenerate separately. ESI does not exist in the free OpenLiteSpeed.

Why cart and checkout are never cached

WooCommerce documentation is explicit: the Cart, My Account and Checkout pages “need to stay dynamic since they display information specific to the current customer and their cart”. The cookies to exclude are named: `woocommerce_cart_hash`, `woocommerce_items_in_cart`, `wp_woocommerce_session_`, `woocommerce_recently_viewed`. The LiteSpeed plugin excludes them automatically in its default configuration.

The circle closes with the Cloudflare `Set-Cookie` rule: those same cookies are why those pages return MISS. It is not a fault. It is the only thing stopping you showing one person’s basket to another.

And a structural gain rather than a hardware one: High-Performance Order Storage moves WooCommerce orders out of the generic content tables into dedicated tables with their own indexes, “which results in fewer read/write operations and fewer busy tables”.

Interior of a modern metro station, with empty escalators and lines of light across the ceiling
Capacity shows when it is empty. A loaded server looks exactly like an idle one, right up to the day with traffic.

Gate three: can anything understand you?

The page is readable and appears quickly. Now comes the part that is no longer mechanical: what it says and how it is organised.

Clear writing is not an aesthetic preference

When a generative system composes an answer, it has to extract from your page a statement it can reproduce without breaking it. A forty-word sentence with three subordinate clauses is hard to extract correctly. A sentence that says one thing, with the subject at the front, is easy.

This is not a new SEO technique. It is good writing, which happens to have become a technical advantage. The working rules we use: one thought per sentence; the technical term explained in the same sentence it first appears; the quotable 25–30 word definition placed immediately under the heading that promises it; the number with its denominator beside it, not three paragraphs away.

Google stated the underlying principle in the helpful content update announcement of 18 August 2022: “content created primarily for search engine traffic is strongly correlated with content that searchers find unsatisfying”. That update introduced a signal at the level of the whole site, not the page — so a section of copy written for robots drags the rest down with it.

On 5 March 2024 Google announced three new spam policies, including scaled content abuse, and estimated the update would reduce low-quality, unoriginal content in results by 40%; on 26 April 2024 it revised that to 45% achieved. The figure is Google’s internal assessment with no published methodology, so it is not independently verifiable — but the direction of the policy is clear and has held.

Structured data: how you tell a machine what you are

Structured data is a hidden block in the page, written in a fixed format, that states explicitly what the information on it represents. For a local business, Google requires only name and address, recommends phone, opening hours and coordinates, and asks for the most specific subtype available: “Use the most specific LocalBusiness sub-type possible; for example, Restaurant, DaySpa, HealthClub, and so on.”

In our set: 119 sites out of 231 carry some structured data block, but only 17 carry a local business one. The most common type found is `WebSite`, then `SearchAction` and `WebPage` — precisely what the theme generates on its own, knowing nothing about the business.

One trap that costs: Google declares ineligible for review stars any page where “the entity that’s being reviewed controls the reviews about itself”. The rule covers embedded third-party review widgets on your own site, not just hand-written markup.

FAQ: useful to people, useful to machines, and a change from May 2026

A question-and-answer section is structurally the easiest format to extract from a page: a question phrased as a question, a short answer attached to it. When someone asks a system exactly that question, you already have the answer written in the shape required.

What changed: FAQ rich results stopped appearing in Google Search on 7 May 2026, and on 15 June 2026 Google removed the feature’s documentation entirely. The deprecation concerns the accordion shown on the results page — it does not ban question-and-answer content on the page, and does not make it useless. It only removes it as ornament in the SERP.

In our set, FAQPage schema exists on 5 sites out of 231. 2.2%.

Images and video

Google documents that it uses alt text “along with computer vision algorithms and the contents of the page” to understand an image’s subject. It recommends short, descriptive filenames — while honestly noting they give “very light clues”, which contradicts the habit of stuffing keywords into filenames.

For video, the condition matters and is routinely missed: “The indexed watch page must be performing well in Search before its video can be considered for indexing.” Video does not rescue a weak page; it amplifies one that already works.

We found no published Google figure on the quantitative effect of images or video on ranking. There are documented best practices, not measured effects. Anyone telling you “video raises your position by X%” invented X.

One thing Google says plainly about AI

In its documentation on AI features, Google writes: “There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary… There’s also no special schema.org structured data that you need to add.” It refers strictly to the features inside Google Search, says nothing about ChatGPT or Perplexity, and does not contradict the need for classic good practice. But it cuts the root out from under an entire genre of commercial offer.

Pale render of a neural network: thin lines connecting small orange points on a white background
What a generative system extracts from a page is a statement, not a page. If the statement will not come away cleanly, it does not come away at all.

Gate four: does anyone believe you?

The last gate is the slowest and the hardest to fake. Everything here is what somebody other than you says about you.

E-E-A-T, exactly what it is and what it is not

E-E-A-T stands for Experience, Expertise, Authoritativeness, Trust, and it is defined in Google’s Search Quality Rater Guidelines — a 182-page document, the version dated 11 September 2025. Of the four, Google names trust the most important: “untrustworthy pages have low E-E-A-T no matter how Experienced, Expert, or Authoritative they may seem”.

And the sentence worth reading before buying “E-E-A-T optimisation”: Google’s public documentation states textually that “While E-E-A-T itself isn’t a specific ranking factor, using a mix of factors that can identify content with good E-E-A-T is useful.” It is not a button. It is not a score. The guidelines also specify that no rating given by a human evaluator can move a page in results.

What remains concrete: who wrote the text, how it was made, why it exists. Google calls “Why” the most important of those three questions. For subjects that can significantly affect people’s health, financial stability or safety — the category called Your Money or Your Life — “our systems give even more weight to content that aligns with strong E-E-A-T”. That bears directly on medical practices, dental practices and law firms.

Google Business Profile

For a business with a physical address, the map profile is often more important than the website. Google publishes the three local ranking factors — relevance, distance, prominence — without publishing the weights. And it states something worth remembering before any commercial conversation: “There’s no way to request or pay for a better local ranking on Google.”

What is up to you, from Google’s own documentation: the real name, used consistently on the storefront, website and stationery — extra information in the name field can lead to suspension of the profile. A precise address; post office boxes at remote locations are not accepted. The most specific category — Google gives the example itself: instead of “Salon”, choose “Nail salon” — and explicitly not a category for every service. Attributes appropriate to the category.

On reviews, Google says “more reviews and positive ratings can help your business’s local ranking”. “Can help”, not “determine”, and with no published threshold.

Reviews, and where the forbidden zone starts

This is where people go wrong most often and most confidently. Google policy explicitly prohibits paid reviews “directly or in kind”, and classifies as fake and misleading content the offering of free or discounted goods or services in exchange for posting, changing or removing a review. The prohibition runs both ways — including paying to remove a negative review. It is also prohibited for a merchant to ask staff for a certain number of reviews, or for reviews containing specific content.

Neutral solicitation, with no incentive and no imposed content, remains permitted by the same policy. In practice: “an honest opinion would help us” is fine, “a free coffee for five stars” is not.

The sanctions are real: Google can block new reviews, unpublish existing ones, and display a public warning on the profile that fake reviews were removed.

On third-party platforms, the scale of the problem is published. Trustpilot removed 4.5 million fake reviews in 2024 — 7.4% of everything submitted that year — and 90% of them were removed automatically. Tripadvisor received 31.1 million reviews in 2024 and blocked 2.7 million as fraudulent, also removing 214,000 AI-generated reviews. The figure that should give anyone in hospitality pause: on Tripadvisor, “review boosting” — positive reviews posted by owners, staff or affiliates — accounted for 54% of all detected fraud. The commonest form of fraud comes from businesses, not from competitors.

Trustpilot in turn prohibits incentivised reviews, declaring ineligible any reviewer who has received or been offered an incentive — discounts, promo codes, prize draw entries, refunds, free products.

Links, in 2026

Google defines link spam as creating links “primarily for the purpose of manipulating search rankings”. The criterion is intent, not the existence of the link. The policy explicitly lists as violations: exchanging money for links or for posts containing links, exchanging goods or services for links, and sending someone a product in exchange for an article with a link. That last one catches influencer collaborations and local barter, which nobody thinks of as buying links.

The clean path is written in the same place: it is not a violation to have such links “as long as they are qualified with a `rel=nofollow` or `rel=sponsored` attribute value”. You do not hide the relationship, you mark it in code.

How much links weigh today is a question Google answers cautiously: it uses them “as a signal when determining the relevancy of pages and to find new pages to crawl” — a signal. Gary Illyes, from the Search team, said publicly on 21 September 2023 that links are no longer in the top three factors “and it hasn’t been for some time” — a three-year-old statement, reported by the trade press rather than an official announcement. An indirect clue: the last update explicitly named “link spam update” is the one from December 2022; since then the work has been folded into general spam updates.

Directories, consistency, branding

Here we have to be honest about what is not documented. The idea that consistent name, address and phone across third-party directories is a ranking factor has circulated in the industry for a decade. We found no Google documentation asserting it. The only consistency requirement Google writes down concerns the business’s own materials — storefront, website, stationery — not directories.

That does not make directories useless. It means the argument for them is a different one: they are places a person can find you, and sources a generative system can read. A commercial study across 28 million queries counted 512,680 Yelp citations in AI systems in the fourth quarter of 2025, 3.4 times more than the next platform — but it is an agency study, on the US market, where Yelp is large and in Romania marginal. We cite it as a clue about mechanism, not as a transferable number.

Branding enters here through the back door, and it is the hardest thing to sell as a service, because it has no button. If people search for your name, you have a signal nobody can buy. Social presence works the same way: not as a direct factor, but as the place where name searches happen and where the reason to search for it gets built.

A closed padlock on a metal fence at dusk, with blurred lights behind it
Trust is the one thing in this article nobody can install on your behalf.

The process, from zero to steady growth

Everything above is parts. The order in which you touch them matters more than any single one, because some have no effect at all until the ones before them are done. There is no point writing thirty articles for a site Google cannot read.

Before the steps, a statement that is not rhetoric: everything that follows is work done in the open. Nothing against Google’s rules, no shortcuts, no network of sites built to link to each other, no mass-generated copy, no arranged reviews, no page served differently to a robot than to a person. Not out of delicacy — because all of those have an expiry date, and when they expire you lose exactly the asset you paid to build. Where a tactic is tempting and risky, we name it and say why we do not use it.

Step 0 — The dated baseline

Before any change: Search Console verified, analytics installed, an export of today’s state. How many pages are indexed, which queries bring impressions, what the average position is, how long the page takes. Without that photograph, in a year nobody can demonstrate anything — not you, and not whoever works with you.

This is also where an administrative point gets settled that turns out to be decisive at separation: the accounts are yours. Search Console, analytics, tag manager, the map profile, the domain, the hosting — all on your email address, with collaborators added as users. If the agency owns them, the day you part ways you lose the entire history.

Step 1 — Brand level: who you are and why it would matter

This sounds like a creative workshop and is not. It is a list of questions with written answers, because each of them later becomes text on a page: what exactly you sell, to whom, what makes you different from the business next door in checkable terms, what you do not do, who the people behind it are and why anyone would trust them. Google structures content evaluation around Who, How and Why — and calls “Why” the most important. If the answer to “why does this site exist” is “so it shows up in Google”, you also have the answer to why it does not.

Step 2 — The technical foundation

In the order we walked through above, and in the order it gets repaired:

  1. Check what the site answers to a robot, not to a browser. If it returns 403 or 429 to legitimate requests, that is the hosting security module, not the site.
  2. Write `robots.txt` as a decision, not an inheritance. Separate training from real-time citation. Check what you have now, not what you think you have.
  3. A correct sitemap, with a real `lastmod`. If your system stamps the build date everywhere, it is better to drop it.
  4. Verify the property in Search Console and read the indexing states. Fix “unknown” first — that is a discovery problem — then the rest.
  5. Speed: measure twice, with a control. Fix what is on the server before touching images.
  6. Local business structured data, with the specific subtype. Two mandatory fields.

This step ends. It is not “ongoing optimisation”, it is a repair with a finish line. If someone invoices you monthly for “technical optimisation” indefinitely, ask what gets repaired in the month you are paying for and how it is verified.

Step 3 — Planning: which questions deserve answers

This decides a year of work. The method, in order: what people already search for that reaches you (from Search Console, if you have history); what customers ask on the phone, which is almost never what the site says; which answers are weak or missing on the pages currently ranking.

The output is a list of subjects with one intent each. One subject, one page. Two pages answering the same question fight each other and both lose — the commonest way a site with plenty of content has very few results.

Step 4 — Service pages before the blog

This order gets reversed often, and wrongly. The pages describing what you sell are the ones that must win commercial searches. Articles work for them, not the other way round. If your article outranks your own service page on the same search, you do not have a good article, you have a weak service page.

Step 5 — The editorial plan, and why cadence is not a virtue

We learned this the hard way in the summer of 2026. We were publishing daily, in three languages, and at some point we noticed results sliding. The diagnosis was useful precisely because it was not technical: the same pages were being picked up without trouble by other engines over the same period. Nothing on the site had broken — we had simply outrun the rate Google was willing to ingest. We corrected by cutting the cadence to one article a week.

The practical conclusion for anyone: the rhythm that matters is the one you can sustain at constant quality, not the largest one you can produce. One piece a month that answers a real question completely beats four a month that touch it.

Step 6 — Links, the slow clean part

There is no fast, safe method. There are slow safe methods, and fast methods that work until they stop working. What we do: content someone has a reason to cite, because it contains a measurement or a fact that does not exist elsewhere; real partnerships with suppliers, customers and associations, where the link is a consequence of the relationship rather than its object; presence in directories relevant to the trade; showing up to what happens locally.

What we do not do, and why: private blog networks, because that is the textbook definition of link spam; buying links without marking them, because it is listed explicitly as a violation; reciprocal exchanges at scale; comment and profile links. If you receive an offer of “N links a month at a fixed price”, you are looking at exactly the practice Google names.

Step 7 — Ongoing optimisation, which means something other than you think

Once the foundation is done and the first pieces are published, the monthly work is no longer “more pages”. It is: reading new queries in Search Console and writing answers for the ones that appear; strengthening pages sitting at positions 11 to 30, where the smallest effort yields the most; updating copy that has aged; checking nothing broke technically after a theme update.

That part has no end, because the competition has no end either. But it is measurable month by month, and that is the test. This is what our SEO and AI search work looks like once the repairs are finished: a report that says what was touched, what moved and what did not.

Empty industrial space with a high ceiling and a long row of windows showing the city in the distance
Walls before furniture. The order looks obvious until somebody sells you content for a site Google cannot read.

The tools and the decision each one informs

A list of tools without explanation is a brochure. Each of the following has a job and a decision it informs; if you do not use the decision, you do not need the tool.

What each tool measures and what you decide with it
ToolWhat you seeThe decision
Google Search Consolewhat is indexed, which queries you appear on, at what position, with what click-through ratewhich page deserves strengthening and what copy needs writing. The only source for “am I visible?”
Bing Webmaster Toolsthe same, on the other half of searchwhether an indexing problem is yours or Google’s
Analyticswhat people do after they arrive: where from, how long they stay, what they do nextwhether the traffic that is growing is also the traffic that buys
Tag managernothing on its own — it is where measurement codes are administeredonly if you have more than two or three codes. Otherwise it adds weight for nothing
Hotjar or equivalentwhere the finger stops on the page, where the form gets abandonedwhat to change on a page that has traffic but no enquiries
PageSpeed Insightsspeed, lab and fieldwhether the problem is the server, the theme or the images. Two runs, not one
Google Business Profilethe searches that reach the profile, calls, route requestswhether the map deserves investment before the site

A note about tag managers, since they get installed reflexively: they cost you speed. If all you have to measure is analytics, installing it directly is faster. The decision is “do I need to change measurement codes often without publishing the site?” — if the answer is no, you do not need one.

Tall pole carrying several floodlights, photographed from below against a blue sky with scattered cloud
Each tool lights a different patch. If you never look at what it lights, you have bought a pole.

Bing, and the half nobody measures

Almost everyone looks only at Google, and in Romania that is a rational choice, because Google holds by far the largest share of searches. But for two years Bing has become interesting for a new reason: it is the index behind Copilot. Whoever is invisible in Bing is invisible there too. ChatGPT is often said to run on Bing as well, and it did at the start. Since then OpenAI has built its own index and fills it with `OAI-SearchBot`, the crawler measured above. They are two separate doors, not one.

Here we can use our own data, because it is our own site and it is not about money.

05 · Our own data

Bing went from 675 to 928 indexed pages in five months

The number of cittago.com pages in the Bing index, read from Bing Webmaster Tools. The growth does not come from a campaign: it comes from publishing and from submitting new addresses automatically.

+253pages in five months
3,312links in the Bing index
02505007501,000Apr675May735Jun734Aug888Sep928

Source: Bing Webmaster Tools for cittago.com, read on 11 September 2026 through the API. Cumulative snapshots, not summed. cittago.com

Bing went from 675 to 928 indexed pages in five months
value
Apr675
May735
Jun734
Aug888
Sep928
April–September 2026. There is no July snapshot in the account, which is why the line skips the month.

The number of cittago.com pages in the Bing index rose from 675 on 6 April 2026 to 928 on 10 September 2026. Links in the index, from 1,718 to 3,312. The growth does not come from a campaign and does not come from a budget: it comes from publishing and from submitting new addresses automatically at the moment of publication.

The contrast with Google over the same window is what made us change cadence. Same pages, same publication day: Bing took all of them, Google left them waiting. When two engines behave that differently with identical content, the content is not the problem.

And the part that surprised us most, because it is exactly this article’s subject: our positions in Bing are far better than in Google on commercial queries about Cluj. `digital advertising agency` — position 1. `sme advertising 2026 europe` — position 1. `digital solutions cluj napoca contact` — position 5. `al agencies cluj-napoca` — position 6. `agentie digitala` — position 6. `creative cluj napoca service` — position 8. On Google, the same commercial intents find us at positions 28 to 40.

Now the part we have to say so you do not build a plan on an illusion: Bing’s share in Romania is small. No first position changes that, and nobody builds their growth there.

So the correct conclusion is not “move the budget to Bing”. It is this: if Google keeps you waiting for months, Bing tells you within days whether the problem is yours or theirs. It is the cheapest diagnostic instrument you can have, it is free, and it is also the door to the generative systems that do not look at Google. Worth registering, not worth a budget.

One methodological limit, so nobody else has to discover it: keyword research in Bing Webmaster Tools does not work usefully for Romanian — nearly every term tested returns zero, because the volume is too small. For Italian it works. So do not use Bing as a volume source for the Romanian market.

Telecommunications tower with antennas and a white dome, photographed from below against a clear blue sky
The second signal is not stronger. It is simply a different one, and it reaches places the first does not.

What changes from one trade to another

The process above is the same for everyone. What differs is where the fight happens and what you are allowed to write.

Restaurants, cafés, bars

The fight happens more on the map than on the site. The specific profile category matters, and Google has a dedicated menu editor for food and drink businesses, noting that changes can take 24 to 48 hours to appear in Maps and Search. The profile menu is a separate surface from structured data on the site — we found no Google documentation for a dedicated menu rich result in classic organic search.

On the site: `Restaurant` structured data, correct opening hours, a menu in text — not in an image or a PDF, because nothing reads it there. In our set, 4.1% of restaurants had local business schema, and 42.5% blocked GPTBot.

Hotels

The most complicated category, and in our measurement the one that blocks most — 52.2%. Google offers hotels free booking links, officially described as “unpaid links, ranked according to their utility to users”, as opposed to paid hotel ads. Access requires a Hotel Center account and a technical integration, so it is not a surface a small hotel occupies simply by writing well.

Cabins, guesthouses

There is a concrete trap here. `VacationRental` structured data looks made for you, but Google limits it to sites meeting eligibility criteria that already have a Google technical account manager and Hotel Center access. It is not open to anyone who implements it, unlike `LocalBusiness`. So stay with `LocalBusiness` and the right subtype, plus the map profile and the booking platforms.

In our set, cabins and guesthouses had the lowest local schema rate — 3.2% — and blocked GPTBot in 45.2% of cases.

Dental practices

Romania’s dental deontological code defines advertising as any presentation of the dentist’s activity and services for the purpose of promoting them to the public. It prohibits financial inducements and discounts, with an express exception: prices may be displayed on the website. It prohibits claims about the quantity of results or success rates, except where these can be documented. And it explicitly prohibits “garantarea directă sau indirectă a unui act medical stomatologic sau a unui rezultat predefinit” — guaranteeing, directly or indirectly, a dental procedure or a predefined outcome.

Interestingly, in our measurement dental practices came out best on the technical side — 83.3% with structured data, 91.7% with a sitemap, and the lowest AI-bot blocking rate at 25%. The sample is twelve sites, so read it as an indication.

Medical practices

Romania’s new medical deontological code was adopted on 30 October 2025, published in the Official Gazette on 10 December 2025, and came into force on 1 January 2026, repealing the 2012 code. Much of the material online still cites the old version.

It states that advertising of medical services “are rol exclusiv informativ” — has an exclusively informative role — and limits permitted content to a closed list of ten categories of information. A website and a social page “pot avea ca scop doar informarea publicului în legătură cu activitatea profesională”. Comparative, disparaging or superlative claims are prohibited, and a doctor may not associate their image with price reductions or other material advantages — unlike dentists, where displaying prices on the site is expressly permitted.

We give no legal verdicts and do not interpret. We report what the text says, with the quote. For application to your case, ask the college.

Law firms

Romania’s bar council decision 195/2021, in force since 1 October 2021, repealed the old articles 244–250 of the profession’s statute entirely and moved the substantive rules into annexes. Many commercial legal databases still display the old form.

What is explicitly permitted, from Annex XXXIV: listings in directories and online platforms, including social networks; an internet domain, blogs and own pages on social networks; and — the most direct wording — “Paginile de prezentare și orice altă formă de publicitate în rețelele sociale și în motoarele de căutare online sunt permise”: presentation pages and any other form of advertising on social networks and in online search engines are permitted. It is the only explicit mention of search engines in the Romanian regulation of lawyers, and it does not distinguish between organic and paid results.

What is prohibited: naming clients from the portfolio and identifying the litigation the firm has been involved in. But the same guide separately notes that presenting an objective successfully achieved for a client is not a breach — so anonymised case studies are not prohibited as such. There is also a rule that bears directly on SEO: the domain name “nu poate fi formată exclusiv din termeni generici cu referire la serviciile avocatului” — it cannot consist exclusively of generic terms referring to the lawyer’s services. An exact-match keyword domain falls outside the rule.

One point that applies to all three regulated professions: the rules above are about advertising. They do not stop a practice or a firm from having a fast, readable site with correct opening hours and answers to the questions people actually ask. The organic work described in this article sits almost entirely on informative ground. And for these trades, where Google applies extra weight for subjects that can affect health or financial stability, the quality and verifiability of the text matters more than anywhere else.

Silhouettes of high-voltage pylons and thin branches against an orange sky at sunset
Same network, different branches. The process does not change from one trade to another; where the fight happens does.

Why organic looks different when sales are falling

We have left the commercial argument until last, on purpose, because it is the part that is easiest to get wrong by starting from it.

In July 2026 the volume of retail sales in Romania was 5.7% below July 2025. In the same month, the European Union grew 1.0%. It is not one bad month: from November 2025 to July 2026, every month was negative in Romania and positive in the EU27. The consumer confidence indicator stood at −32.6 in August 2026, against −15.0 for the European average.

Those figures are Eurostat, on the deflated series — so the fall is not a price effect, it is less merchandise sold. The European Commission’s spring 2026 forecast put Romanian GDP growth at 0.1% for this year, with recovery only in 2027, and wrote explicitly that “fiscal consolidation and high energy price inflation are likely to further depress real disposable income”.

What a business owner does when they see that in their own takings is predictable: they cut the first line that can be cut immediately. Usually, that is advertising.

With one of our Romanian clients we did the exact opposite, in June 2026. Demand for their services had dropped hard, and our proposal was to double the Google Ads budget rather than cut it. They accepted on trust — we could not guarantee them anything, and we did not pretend to. What happened: they felt a sharper demand for services and purchases almost immediately, their traffic doubled, and sales nearly doubled with it. From next month the budget goes up another 33% on top of the current one.

The part that matters here is a different one, and it is why this story opens an article about organic: in the two to three months since the increase, their Google Business Profile reviews nearly doubled, and their organic clicks and traffic grew too. We paid for none of the organic side. It came because more people reached them, bought, came back and wrote about it. Paid and unpaid are not two separate buckets.

One case, not a rule. We are not saying that any company doubling its budget doubles its sales — we do not know that and nobody does. We are saying that the reflex cut is not automatically the right decision, and that in a market where everyone else is cutting, the auction gets emptier for whoever does not.

Which is where organic becomes interesting, for one reason only but a large one — it has no off switch. An article bringing enquiries in March brings them in November too, without anyone pressing anything.

It also has a disadvantage worth writing in the same sentence, so it is not a surprise six months later: it does not come fast and it does not come guaranteed. If you need customers right away, this is not where they come from.

And it is not free, though that is what it gets called. Position in results is not paid for at Google and not paid for at ChatGPT — there is no button to buy. But the time of whoever writes, the tools that measure, the hosting that keeps the site fast, and the hours spent fixing what is broken all cost. When someone says “free traffic”, translate it as “traffic you pay for once, in work, rather than every time, in clicks”.

One more thing the European data shows, which changes the calculation for a Romanian business. In 2025, 53.6% of Romanian enterprises with at least ten employees had a website, against 79.02% across the EU27 — a gap of 25.4 percentage points. Among small firms of 10 to 49 people, the ratio was 51.02% against 76.67%. In other words: among firms with at least ten employees, your local competition is, statistically, less present online than the equivalent competition elsewhere in Europe. That gap is bad news for the economy and good news for you, if you are willing to do what your neighbour does not.

An important caveat on the Eurostat figures, because you will see them quoted everywhere without it: the survey covers only enterprises with at least ten persons employed. Microenterprises under ten — which is most practices, cafés and guesthouses in Romania — are not in the sample at all. So nobody knows the real gap among the very smallest firms.

White render of a tower built from small stacked cubes, on a light background
It accumulates piece by piece and shows nothing day to day. Which is why nobody honest can promise you the day it finishes.

What you can ask for, and what nobody can guarantee

Anyone who guarantees you a result is the first one to turn down. Not necessarily because they are dishonest, but because nobody controls someone else’s ranking. A position guarantee is a promise made with another party’s property. Google writes it itself about the local side: “There’s no way to request or pay for a better local ranking on Google.”

That does not leave you without criteria. It leaves you with better ones, because they are checkable. Here is what we consider genuine signals of trust.

  1. Real case studies, with reviews from the clients in those case studies. Not a wall of logos. A case described, with what was done, and a person confirming it.
  2. Track record. It does not guarantee competence, but it shows the firm has lived through more than one algorithm change.
  3. Monitoring and honest reports, at a fixed cadence. A report that includes what did not work is worth more than one containing only increases.
  4. You see the results yourself, directly in Search Console and analytics. Not in a PDF someone else built.
  5. They explain the tools enough for you to check in real time. An agency that does not want you to learn to read Search Console has a reason.
  6. The accounts are yours. Search Console, analytics, tag manager, map profile, domain, hosting — on your address, the agency added as a user.
  7. A dated baseline, taken before the first change. Ask for it in writing in the first week.
  8. They tell you what method they use for links. “We have a network of sites” or “N links a month at a fixed price” means buying links, which is exactly what Google lists as a violation.
  9. They tell you who writes the copy. “We generate it with AI at volume” is a policy risk, not an efficiency — scaled content abuse has been an anti-spam policy since 5 March 2024.
  10. A change log. What changed, when, why. Without it you cannot attach a rise to a cause.
  11. They tell you what stays yours if you leave. The copy, the structure, the structured data stay, if they were built on your site. They do not, if they were built on the agency’s subdomain.
  12. They give you a realistic range and tell you what shows up first.

What to watch, specifically, in Search Console and analytics

The signals appear in an order, and that order is the best theatre detector you have:

The order in which the signals of real organic growth appear
OrderWhat movesWhere you see itHow easy to fake
1new impressionsSearch Console, Performancehard
2new queries in the listSearch Console, Querieshard
3click-through rateSearch Consolemedium
4average positionSearch Consoleeasy, by choosing the word
5clicksSearch Console and analyticsmedium
6enquiries, orders, callsanalyticshard

From analytics, what matters in addition: traffic source — is organic growing, or just the total? — bounce rate, and average time on page, rising or falling. Traffic that grows while bounce rate grows alongside it means you are attracting the wrong people.

And the most useful point in this whole section: an agency that shows you “position 1” first, on a word nobody searches for, has skipped the first four signals. Impressions and new queries appear first and are the hardest to fake. Positions appear last and are the easiest to cherry-pick.

What is not a signal: the number of “optimised keywords” · third-party tool scores presented as a promise · “we guarantee top 3” · reports with positions but no impressions and clicks beside them · traffic growth with no breakdown by source.

If you want to see what this work looks like on our side, the SEO and AI search page says what we do and what we do not.

Blurred city lights seen at night through a windscreen covered in raindrops
From the threshold you can see which way, not how far. That is all a measurement taken today can honestly say.

Three thresholds, so you know where you start

Not everyone is in the same place, and not everyone needs everything above. Three situations, and what to do in each.

If your site answers a robot with 403 or 429, or if Search Console shows pages as “URL is unknown to Google”: you are before the starting line. Nothing you write matters yet. This is a few hours of work with whoever runs your hosting, and it is the cheapest intervention in the whole article. Check today.

If the site is readable but has no h1, no description, no local business structured data: you are at the starting line with the wheels off. That is a day or two of work, measurable within weeks, and it needs no monthly budget. A third of the sites we measured are here.

If the technical side is sound and you already appear at positions 11 to 30 on queries that matter: you are in the race, and from here it is long work — content, authority, patience. This is where paying someone makes sense, and where the checklist above applies.

The measurement in this article is from 11 September 2026 and it is a photograph, not a film. We may repeat it in 3–6 months with the same set rule, to see whether AI-bot blocking rises or falls — and in that case we will try to publish the results. Let’s talk about this again in 3–6 months 😉

A thin thread of white smoke curling against a black background
The questions left once the numbers have settled are usually the useful ones.

Questions with answers we think are useful

What does it mean that my site “blocks GPTBot”?
It means your server answers a request introducing itself as GPTBot differently from one introducing itself as a browser — usually with an error code, 403 or 429. In our measurement on 11 September 2026, 97 of 228 Cluj sites did this, and only 11 had written it in robots.txt. The rest inherited it from their hosting security module.
Is it bad to block AI crawlers?
Not necessarily. It is a legitimate business decision in both directions: a publisher living on advertising has reasons to refuse training, and a restaurant that wants to be recommended when someone asks an assistant where to eat has reasons to accept. The problem is deciding nothing and not knowing what you do. Cloudflare separates the two cases into distinct signals: `ai-train` for training and `ai-input` for real-time use inside an answer.
How do I check this myself, without paid tools?
Open `yourdomain.com/robots.txt` in a browser and read it. Then verify the property in Google Search Console — it is free — and look at the indexing report. To test bot blocking you need someone who can make a request with a changed user-agent; it cannot be done from a browser.
Does the measurement apply outside Cluj?
We do not know, and that is the honest answer. The set is Cluj county, defined by an OpenStreetMap rule, and it is a convenience sample rather than a representative one. What we would expect to travel is the mechanism, not the percentages: hosting security modules ship with the same defaults across markets, and nothing about them is specific to Romania.
Does llms.txt matter?
We do not know, and it is right to say so. OpenAI’s crawler documentation does not mention it at all. Anthropic publishes its own files, but publishing does not prove consumption. Google says in a guide updated on 10 July 2026 that you do not need AI text files to appear in Search. It is cheap to add and has no known downside — we have one — but it is not a lever.
How long until I see results from SEO?
It depends where you start, and anyone giving you a fixed number without having seen your site is guessing. What we can say with confidence is the order: new impressions and new queries appear first in Search Console, clicks and positions come after. If in three months not even the number of queries you appear on has moved, something is wrong with the approach.
Can I ask customers for reviews?
Yes, if you offer nothing in return and do not tell them what to write. Google policy prohibits reviews paid for “directly or in kind” and treats offering free or discounted goods or services in exchange for a review as fake and misleading content. The prohibition runs the other way too: you cannot pay to have a negative review removed. You also cannot ask staff for a certain number of reviews.
Why does PageSpeed give me a different score every run?
Because the lab test runs with simulated throttling — 1.6 Mbps, 150 ms latency, a processor slowed fourfold — and Google explicitly lists the sources of variability: local network, hardware and resource contention. Which is why our rule is two runs plus a control. Field data also moves with a delay: the collection window is 28 days.
What is the difference between GPTBot and OAI-SearchBot?
GPTBot collects pages to train OpenAI models. `OAI-SearchBot` is the crawler behind search inside ChatGPT, which is what happens when the assistant looks something up on the web to answer you. They are two separate crawlers with two separate permissions: you can refuse the first and accept the second, so that you are not used for training but can still be cited. In our measurement on 11 September 2026, out of 225 Cluj sites, 94 refused GPTBot and 71 refused `OAI-SearchBot`; 26 refused training only.
Where do the numbers in this article come from?
The ones about the Cluj market were measured by us on 11 September 2026, on a set defined by rule from OpenStreetMap, and contain no domain names. The ones about cittago.com come from Google Search Console and Bing Webmaster Tools for our own site, and concern indexing and positions only, not money. The rest are from vendor documentation and official sources, each linked at the end. No figure comes from any client’s advertising account.

Last updated: 11 September 2026. The Cluj market measurement was run the same day, on a set defined by rule from OpenStreetMap (Cluj county, relation 91733, seven categories, units with a website filled in on the map): 1,782 mapped units, 314 unique domains, 231 that answered, 228 in the user-agent test, 225 in the second round that added OAI-SearchBot. The cittago.com figures come from Google Search Console and Bing Webmaster Tools, read the same day through their APIs, and concern indexing and positions only. Every other figure comes from vendor documentation or official sources, each opened and cited on 11 September 2026. No figure in this article comes from an advertising account.

Sources: Google · Common crawlers · OpenAI · Overview of OpenAI Crawlers · Anthropic · crawlers and robots.txt · Perplexity · PerplexityBot · Meta · Web Crawlers · Cloudflare · Content Signals Policy · Cloudflare · AI traffic options, 1 July 2026 · OWASP CRS · Anomaly Scoring · Google · Page Indexing report · Google · URL Inspection Tool · sitemaps.org · protocol · llmstxt.org · the llms.txt specification · Google · guide to generative features · Google · AI Features and Your Website · web.dev · Core Web Vitals · web.dev · INP becomes a Core Web Vital · Google · About PageSpeed Insights · Chrome · Lighthouse performance scoring · Google, Deloitte Digital · Milliseconds Make Millions · web.dev · Vodafone case study · web.dev · Rakuten 24 case study · HTTP Archive · Web Almanac 2025, Ecommerce · HTTP Archive · Web Almanac 2025, CMS · Cloudflare · default cache behavior · Cloudflare · Automatic Platform Optimization · Cloudflare · API deprecations · WooCommerce · configuring caching plugins · WooCommerce · server requirements · LiteSpeed · LSCache documentation · WordPress · WP_Object_Cache · Google Cloud · disk performance · Google · improve your local ranking · Google · guidelines for representing your business · Google · user generated content policy · Google · Search Quality Rater Guidelines, 11 September 2025 · Google · creating helpful content · Google · spam policies · Google · LocalBusiness structured data · Google · review snippet structured data · Google · Search documentation changelog · Trustpilot · Trust Report 2025 · Tripadvisor · 2025 Transparency Report · arXiv · GEO: Generative Engine Optimization · Answer.AI · the llms.txt proposal · Google · first generative capabilities in Search, 10 May 2023 · Google · AI Overviews rollout, 14 May 2024 · OpenAI · Introducing ChatGPT search · Pew Research Center · clicks and AI summaries · Ahrefs · AI Overviews and CTR, February 2026 · Google · Liz Reid’s response, 6 August 2025 · Google · helpful content update, 18 August 2022 · Google · March 2024 core update and new spam policies · Google · crawl budget guide · Eurostat · retail sales, monthly data · Eurostat · enterprises with a website · European Commission · economic forecast for Romania · Romanian College of Dentists · advertising rules from 1 July 2025 · Romanian medical deontological code, Official Gazette 1141/2025 · Romanian Bar Council · decision 195/2021 and annexes

SEO and AI search

Want to know what your site answers a robot?

We will check the four gates from this article on your domain, free, and send you what we found — with what takes hours to fix and what takes months.

What clients say

Trusted by the people who signed the checks.

5.0★★★★★21 reviews on Google
★★★★★
We have been collaborating for over 11 years on both presentation web sites and complex projects. We have always returned to the services offered by Cittago, thanks to the professionalism, courtesy and innovative solutions offered. Thanks for your partnership!
Aurelia Campean2 years ago
★★★★★
I am very satisfied with the collaboration with Cittago. Everything went in a professional manner, the deadlines were met, and the result was as expected. I highly recommend!
Cristea Christian4 days ago
★★★★★
5* for the quality of service, promptness and seriousness. Thank you, Paul!
Budurlean Crina5 days ago
★★★★★
We had the cabins and the view, but Cittago gave us the perfect digital “reception”. They created a premium, ultra-fast website for us that handles everything on its own: live calendar, automatic invoicing and card payments (only 1% commission instead of 15–20% on platforms). The best part? We edit it ourselves in a few minutes, without depending on anyone. And Paul is simply unreal for this world! The warmth, respect and attention to detail with which he explains absolutely everything make you understand the services offered perfectly. Not to be missed is the availability that Paul shows when you have a question. Honestly, I have rarely dealt with such a professional and dedicated company.
Viorica Pop5 days ago
★★★★★
Serious and fast team. They built our BarBox website from scratch, with a cinematic look that represents us perfectly, plus local SEO so people in Cluj can find us. Simple communication, zero hassle. 5 well-deserved stars.
tudor j6 days ago
★★★★★
I had the pleasure of working with Cittago for the creation of a website for a project that I develope together with some friends, and I couldn't be happier with the results. From start to finish, they demonstrated exceptional professionalism, creativity, and technical expertise. First and foremost, the communication throughout the project was outstanding. They took the time to listen to our ideas and goals, and they translated them into a visually stunning and highly functional website that perfectly represents our product. We were kept in the loop at every stage of development, and they was always quick to address any questions or concerns I had. What sets Cittago apart is their dedication to delivering results. They went above and beyond to ensure that our website met all our requirements and objectives. They even provided valuable suggestions and insights that improved the overall project. I wholeheartedly recommend Cittago to anyone looking for a digital agency that combines creativity, technical expertise, and exceptional customer service. Thank you Paul for a job well done!
Bochiş Răzvan2 years ago
★★★★★
Thank you Paul for all the professionalism you show, for all the patience and all the help you give me. I highly recommend!
Daniela Pasc2 years ago
★★★★★
The collaboration I have had since the beginning, that is, for several years, with Cittago is a real pleasure! I turned to Paul to rebuild the website of a small dental clinic and I am extremely satisfied with the collaboration. Paul is still taking care of the website. The promptness with which he responds to me, the patience with which he explains everything I don't understand (and there are many, believe me 😂🙈), his maximum involvement and desire to give his best have always helped me and given me a lot of confidence in him. He is always there when I need him. Very professional! And the quality-price ratio is unbeatable. I recommend Paul with confidence, if you want someone who really puts his heart into what he does and gives his best!
Daniela Chis2 years ago
★★★★★
I have had the pleasure of working with Paul and Cittago on several websites. From the initial discussion to the launch of the website, I was impressed by their professionalism, expertise, and dedication to creating an exceptional product to launch online that we can all be proud of. Paul took the time to truly understand what I wanted, what my brand meant, what my target audience was, and what my goals were. Using his knowledge, he was able to create a user-friendly, visually beautiful, responsive (on both mobile and desktop) website that communicates the services and products we offer very well. The website not only looks great, but it works just as well. Throughout the process, Paul was receptive to feedback, and patiently answered all the requests I had. We appreciated his vast knowledge (about website creation, SEO, social media connection, visual experience), his expertise in applying it, his patience, transparency, and his ability to successfully complete such a project, which was very important to us. For these reasons, I highly recommend this company and the services it offers.
Aissa Suciu2 years ago
·First paint — when something appeared·Server response — before anything could load·Page ready — when you could interact
Page loaded in ·