From the outside, a page no search engine has found is indistinguishable from a page that was never written. This article pulls apart the three things most teams compress into the word "indexing", sets out which of them the Semalt panel's Indexing Hub can actually observe, and names the guarantee that no product in this category is able to make.
Most Calgary firms that call about search do not have a volume problem. A mid-sized engineering consultancy downtown might publish six hundred pages, half written before the current marketing lead was hired. Yet its two most important pages — the practice-area page that wins RFPs, and the profile of the principal clients search by name — took months to appear, if they appeared at all.
The instinct is to blame the writing. Usually the writing is fine. What went wrong sits earlier, where an engine decides whether an address is worth knowing about, worth fetching, and worth keeping.
A page that exists is not a page that is found
Publishing changes the state of your CMS. It changes nothing at Google until a crawler learns the address, spends a request on it, and a separate process decides the result is worth storing. Days, weeks or silence separate those events, and a failure at the first cannot be repaired at the third.
Keeping them apart is not pedantry; it decides which team gets the work. Discovery belongs to navigation and the sitemap, crawling to whoever owns the servers and redirect rules, acceptance to whoever writes the page. Send the wrong problem to the wrong team and it returns unsolved.
| Stage | What the engine is deciding | How it stalls at a head-office firm | Who can actually fix it |
|---|---|---|---|
| Discovery | Whether this address is worth knowing about at all | A practice page reachable only from a dropdown; a sitemap last configured during the 2019 rebuild | Whoever owns navigation and the sitemap file |
| Crawling | Whether to spend a request fetching it now | Requests absorbed by expired postings, print variants and a decade of archives | Hosting, redirect rules, robots and parameter handling |
| Indexing | Whether what came back deserves storage | Forty profiles from one template; three near-identical capability statements | Whoever writes and approves the page copy |
Crawl budget is the second of those three, and it is not something you buy. It emerges from how quickly your server answers, how reliably it answered before, and how much the engine thinks the domain is worth revisiting. A site that stalls under load gets crawled cautiously: the crawler throttles itself rather than knock you over.
Not large, but layered — by history and by function
The sites that give us the most trouble here are not catalogues. They are professional-services firms, engineering and environmental consultancies, and B2B companies with a national client list run from one Calgary address. One domain serves audiences with nothing in common and carries sediment from every earlier version of itself.
Count the strata. People pages, turning over as staff join and leave. Project and case-study pages, the material that decides shortlists. Insight posts published whenever someone senior has time. Careers postings that expire but never disappear. Investor and news releases going back a decade. And underneath, two or three rebuilds still answering on old addresses.
People pages
A principal's profile is often the highest-value page on the site and the shortest-lived.
- Departed staff still returning 200
- New hires missing for months
Projects and case studies
What a prospective client reads before agreeing to a meeting, and the layer most often buried.
- Reachable only through a filter widget
- No crawlable index of the full set
Expired postings
Roles that closed two years ago and releases nobody has read since publication.
- Volume without value
- Still listed in the sitemap
Residue of past rebuilds
Two or three previous structures still resolving, often on paths nothing links to.
- Duplicate paths to identical content
- Redirect chains three hops deep
Put them together and the characteristic failure appears. Six hundred pages that matter compete with two thousand that do not, and nobody has told the engine which is which. The pages carrying your credibility queue behind expired job ads.
What consumes crawl budget when size is not the issue
On a small site, the budget is rarely exhausted by legitimate pages. It goes on machine-generated variants nobody asked for and nobody has audited since launch.
Archive pagination without end
Twelve years of releases generate author, year, month and tag archives, each paginated.
- 140 releases become thousands of addresses
- Every one answers 200
Filters on the project list
Sector, service line, region and year combine into more URL variants than the firm has ever had clients.
- Near-identical content at each
- No canonical declared
Chained redirects
Each hop is a separate request: three hops across two thousand legacy addresses is six thousand fetches arriving nowhere new.
- Chains grow with each rebuild
- Nobody audits them
Postings that never end
Closed roles left live are recrawled indefinitely, because nothing has told the engine they are finished.
- Often the largest section
- 410 is the honest answer
None of this needs a large site — only a CMS left running for several years, which describes almost every professional-services site in the city. The remedy is deletion and correct status codes, not new writing.
The Indexing Hub, and the four numbers that shape every plan
Inside one panel that also holds campaign automation, Search Console analytics, rank tracking and AI market research, the Indexing Hub is the module dealing with everything before ranking. Its job is to push addresses, read sitemaps and keep a record of both; the module overview of the rebuilt Semalt panel describes the rest.
Learn these four figures before you write a schedule
Every sequencing decision in a migration follows from them, and misreading them is how a two-week promise becomes a two-month one.
- A daily budget of 1,000 URLs per account. Throughput, not credit: unused days do not accumulate, and tomorrow cannot be spent today.
- Batches of up to 10,000 URLs. This measures one hand-off, not daily movement. Hand over the maximum and you have queued ten days of work in a single click.
- Three levels of recursion, 1,000 sitemaps in one job. Give the Hub a file or an address. Where an index file points at further index files, it keeps descending — three levels down, a thousand sitemaps per job.
- Two jobs at a time, a queue of twenty behind them. Sitemap jobs run in pairs; number three onwards waits, and the queue tops out at twenty. Fifty scattered submissions all execute, but nobody can say what is where.
Delivery goes out over IndexNow, the API through which GoogleBot, BingBot and other participating crawlers are told an address is new or has been edited. The usual sequence inverts: rather than waiting for a bot to find its way back to a section it last saw in spring, you raise your hand as the change happens. Delay costs most on a new principal joining, or a project page published the week a proposal goes in.
Sitemaps as an instrument, not a formality
Most firms treat the sitemap as a file a plugin writes and nobody reads. Used properly it is the cheapest diagnostic available: a split, honest sitemap shows which layer of the site is being neglected.
Recursive parsing is where the truth emerges, since the Hub reads what the file says rather than what the team believes. A plugin left alone for four years lists staging paths, every expired posting and both structures the site had before this one.
Six corrections worth making before you submit anything
One afternoon of tidying commonly buys back several weeks of budget that would have gone on addresses you never wanted crawled.
- List only addresses that are their own canonical. A page pointing its canonical elsewhere does not belong in the file; listing it asks the engine to settle a contradiction you created.
- Every entry answers 200. A redirect or a dead page in the file burns a fetch and reads as carelessness. Departed staff and closed roles are the offenders.
- Nothing marked noindex appears. Inviting a bot to fetch what you have already told it to discard is the contradiction we find most often on these sites.
- Split by function, deliberately. Separate files for people, projects, insights, careers and news turn one opaque number into five.
- Timestamps must mean something. When lastmod moves during every overnight build, the crawler learns your dates say nothing and stops consulting them.
- Language versions agree in both directions. An English page pointing at a French counterpart needs the return pointer too, and both must match their canonicals.
The two-job limit matters more than it first appears. A structured set of index sitemaps — one parent, five children by function — processes and reports as a single job. Fifty separate submissions all run eventually, since the queue drains, but you lose any view of what has been processed.
When a firm rebrands or merges and the whole structure changes at once
This city changes nameplates more often than most. Consultancies merge, a national practice absorbs a local one, a firm drops a founder's surname, and the domain changes with it. Every address moves in one evening, on a launch date chosen by people who have never heard the phrase crawl budget.
What follows is predictable. Branded search — for a firm whose buyer is a committee, most of the search that matters — starts returning the old domain, a directory listing, or nothing useful. Every credibility page now sits at an address no engine has been told about.
| Decision at launch | What the crawler encounters | Cost to the daily allowance | Better approach |
|---|---|---|---|
| Everything on the old domain redirects to the new homepage | Thousands of addresses collapsing to one destination | Every legacy URL fetched, none usefully | Map page to equivalent page; use the homepage only where nothing equivalent exists |
| Old domain retired at launch | Nothing resolves; accumulated authority is discarded | Not spent, but not recoverable either | Keep the old domain answering with permanent redirects for at least a year |
| Both brands live at once during the transition | Two complete copies of the same firm | Allowance halved across a duplicate set | One canonical brand from day one; the other redirects, never mirrors |
| Everything submitted at once, unsorted | Practice pages queued behind expired postings | Weeks before commercially critical pages are reached | Submit by measured value: revenue pages first, archive last or never |
None of it can be sequenced until you know which pages earned anything under the old name. Clicks, impressions and position history per page come from the Google Search Console views in the panel, exporting up to ten thousand rows as CSV or JSON. Sort that file by impressions and you have the working order; Stream, the assistant sitting in My SEO, accepts such lists a batch at a time.
Reading the status of a batch without fooling yourself
The middle stage is the one the Hub can genuinely evidence. Every address carries an entry: the bot that called, when, the status returned, and the detail behind any failure. Three live counters run above that — submitted, found, failed.
- Two totals, two meanings. The distance between what you submitted and what was found is the discovery gap. Still wide a week later, and the sitemap or the internal linking is at fault.
- A dated 200 records arrival, never acceptance. It confirms the bot asked and the server replied. Storage is a question the log cannot settle, and forgetting that is how teams declare victory early.
- Failures cluster, and the shape names the cause. Thirty errors confined to one path branch point to routing. The same thirty spread evenly point to server load while the crawler was working.
- Judge by impressions, not by the feeling of progress. Two weeks after a batch, check whether the submitted pages have begun collecting impressions. That is the only evidence the third stage went your way.
Common questions
We submitted our new practice pages a month ago and nothing shows. What went wrong?
Open the visit log. An address with no recorded request never reached a crawler: a discovery failure. An address with recorded 200s and nothing in results was read and rejected: a content failure. From outside the two are indistinguishable.
Why is the daily figure 1,000 when a batch holds 10,000?
They describe different quantities. A batch counts what you can queue in one action; the daily figure counts what leaves the queue per account each day. Hand over ten thousand and you have committed ten days, so plan on the daily figure.
Should we push the whole site through and let the engine decide?
Rarely. A page with twelve months of zero impressions wants consolidating or removing, not pushing. Budget spent there postpones the pages that win work, and it re-offers material the engine has already turned down.
Our consultant profiles all use one template. Is that a duplication problem?
It becomes one when the template is most of the page. A photograph, a title and a boilerplate paragraph give the engine no reason to keep forty variants. Two hundred words of specific material — named projects, sectors, credentials — changes that. For the ten people clients search by name, that is one quarter's work.
How long should we keep the old domain running after a merger?
At least a year, and there is little reason ever to switch it off. Permanent redirects to mapped equivalents cost nothing to maintain, and links from industry directories, association listings and old proposals keep arriving. Retiring the domain at launch discards authority that cannot be rebuilt.
The closing calculation, and where to start on Monday
Take a realistic case. An engineering consultancy of 240 staff merges with a smaller practice and moves to a new domain. The combined live site is roughly 900 pages, behind which sit about 2,600 legacy addresses across two old structures. All of it must be discovered again: 3,500 addresses enter the plan.
At a thousand a day the set clears in four days, which sounds like nothing. The arithmetic that matters is different: of those 3,500 addresses, perhaps 180 decide whether the firm survives a shortlist. Those 180 go on day one, and a cleaned sitemap keeps them out of the queue behind eleven years of postings. Most of the 2,600 legacy addresses should never be submitted at all.
Scale it up and the ratio holds. A merged firm with 40,000 addresses faces forty days at full allowance, which is only a crisis if the pages that earn fees sit on day thirty-eight. Sequencing is the discipline; the daily limit is only the arithmetic.
So start narrow, in this order. Split the sitemap by function and drop everything that does not answer 200. Export the pages already collecting impressions, sort by that column, submit in that order. Record the counters as a baseline and read them weekly. Where the engineering queue is the constraint, campaign automation covers the ground that does not depend on it: AutoSEO, at 149 USD monthly per domain, runs keyword discovery and backlink placement with no call on a release slot, and four to eight weeks is the usual wait for measurable movement. Our technical SEO services page describes where an outside team takes over, and to work through the modules yourself, open the Semalt dashboard and connect your property.
Further reading on technical and branded search for Calgary firms sits on our blog.