Backfilling a 1,927-listing directory from other people's metadata
We filled in images, logos, and descriptions across an entire Hawaii business directory by reading each business's own published og: tags. Site-wide the image coverage went from 175 listings to 689 and 432 listings crossed the quality gate. Here are the numbers, the pipeline, and the four traps that cost the most time.
A directory with empty cards is not a directory. It is a spreadsheet with a domain name.
AlohaCalendar has 1,927 active business listings across the islands — activity operators, restaurants, hotels, shops. For a long stretch most of those listings had a name, an island, a category, and nothing else worth rendering. No photo. No logo. A description on maybe one in five. The pages existed, they were crawlable, and they were thin in the specific way that search engines have gotten very good at recognizing.
The obvious fix is to go get photos. The obvious fix is also thousands of dollars of stock licensing, or a photographer, or an intern with a browser and a year. None of those were happening.
The less obvious fix: every one of those businesses already publishes a photo of itself. It is sitting in the og:image tag on their own homepage, put there deliberately so that their link looks good when someone shares it. Same for og:description. Same, often, for a logo in their schema.org JSON-LD. It is the one piece of metadata on the web that exists specifically to be read by someone else’s software.
So we read it.
The numbers
Two rounds on 2026-08-08 — first the 1,033 activity listings, then the remaining 389 restaurants, hotels, and shops. Site-wide, across all 1,927 active listings:
imageUrl: 175 → 689 listingslogoUrl: 3 → 131 listingsdescription: 367 → 1,075 listings- 432 listings crossed the SEO quality-gate threshold — the internal score that decides whether a listing page is substantial enough to be worth indexing. Activity listings went from 42 over the line to 357; the rest went from 186 to 303.
Logo yield turned out to be strongly category-dependent: 17% for activity operators, 29% for restaurants. Restaurants care more about their brand mark appearing in a share card, which makes sense the moment you say it out loud.
The pipeline is reusable and the run order matters, so it lives on its own with a README that explains the failure modes rather than being buried in a one-off script.
Trap one: the image host allowlist that did not apply
The backfill dropped roughly 330 distinct external image hosts into imageUrl. Next.js has a remotePatterns allowlist in next.config.ts that governs which remote hosts next/image will optimize, and ours allowed exactly three.
By every reasonable reading, this was a blocker. Three hundred and thirty hosts, three allowed, and no appetite for maintaining an allowlist of every hotel CDN in the Pacific.
It was not a blocker, because the listing page renders a plain <img>, not next/image. The allowlist never enters the picture.
That is a happy accident rather than a design decision, and it is now a load-bearing one. Migrating that page to next/image — an entirely sensible-sounding future refactor, the kind that shows up in a “modernize the image pipeline” ticket — would break every backfilled image on the site simultaneously. It is written down for exactly that reason. The most dangerous constraints are the ones you satisfy by accident.
Trap two: the column that was populated but never read
generateMetadata built each listing page’s meta description as tagline || generated fallback. It never read business.description.
While the description column was empty, this was harmless — correct, even. The generated fallback was the only thing available.
The moment the backfill populated 1,075 descriptions, that same code became a bug that emits near-duplicate meta descriptions across roughly 1,100 pages, while the real, distinct, hand-published description sat unused one field away. A data migration turned working code into broken code without touching a line of it.
Fixed in the same pass. The general lesson: when you fill an empty column, grep for every reader of that table, not just the writers. Empty columns hide bad reads.
Trap three: a green deploy is not a live fix
Chasing a Search Console alert the same week — “Missing field location”, up 800% — we found that several routes hand-rolled their own Event JSON-LD instead of calling the shared builder, and two of those copies emitted no location at all when a venue was missing. location is required for Event rich results. Without it the event is ineligible for event cards, which is the entire SEO engine of an events site.
Measured live: 11 of 83 Event items were missing location, worst on one island page at 8 of 20.
The fix was not to route the compact summaries through the full Event builder — those emit ItemList summaries, and full Event objects would have bloated the markup and required about fifteen more database columns. Instead the Place-construction logic was extracted into one exported helper now shared by the full builder and all three summary callers, so it cannot drift a fourth time. Deployed, healthy, zero 5xx.
Then the re-count said the fix had failed.
It had not. The deploy script ends in a cache purge that, called with no arguments, purges a five-URL hot set — deliberately, because purging everything leaves every page cold at once and the resulting server-render herd exhausts the app container’s memory. That has bitten a sibling project hard enough to be a standing rule. Any deploy touching pages outside that hot set keeps serving pre-deploy HTML at the edge.
We were reading a cache with age values between 1,194 and 7,087 seconds against an s-maxage of 300. Something in the zone configuration holds HTML far longer than the origin asks for. Stale for hours, not five minutes.
Re-verified with a cache-buster, then purged the changed URLs explicitly: 11 of 83 → 0 of 126 Event items missing location, across the island pages, the event guides, the Japanese localisation, and the weekend pages. A bonus fell out of it — the summary callers used to emit a lean Place with no address, and address was missing on 72 of 164 items in an earlier report. The shared helper is full-fidelity, so that warning should go quiet too.
The rule that came out of this: a green deploy tells you the origin changed. Only a cache-busted request tells you the internet changed.
Trap four: the failure list that is not a failure list
The backfill produced a byproduct — 74 listings whose website did not fetch. It would be very easy to read that file as a closure list and deactivate 74 businesses.
It is not a closure list. Thirty-four of those are plain HTTP 403 bot protection from large hotel and restaurant brands that are extremely open for business. Most of the DNS failures are hotels that moved off vanity domains onto a corporate parent’s URL structure. It is a URL-refresh backlog wearing a scary hat.
We also learned when to stop. The 403s are TLS fingerprinting of the HTTP client, not IP-based — a residential Hawaii address got the identical 403 — so only a real browser gets through, and several major chains block even that. Worse, the prize is not worth the siege: chain og:image values are nearly always a corporate logo rather than a photograph of the property. Net yield across roughly two dozen attempted sites was four field writes, because most of those rows already had better hand-written descriptions. That effort is now explicitly marked do-not-retry.
One more calibration note, since it cost real trust in a data source: schema.org JSON-LD image is markedly less reliable than og:image. One oceanfront hotel’s JSON-LD image was a photograph of Fiji. og:image gets looked at by a marketing person every time someone shares a link. JSON-LD image gets looked at by nobody.
What is honestly still undone
Roughly 8–10% of the shipped imageUrl values are still logos rather than photographs. The detector that sorts one from the other works on colour statistics — flat fields of brand colour with high edge contrast read as logos — and it plateaus right about there. Getting the rest needs text or OCR detection, because past that point the distinguishing feature genuinely is “there are words in it.”
The 34 stale website URLs are worth re-sourcing by hand. Deduplication surfaced 18 name-matched groups of which only 3 were real duplicates; 14 were companion rows deliberately backing a separate weddings vertical, and merging them would have gutted that directory. The three real ones were retired with redirects and a scraper lock rather than deletion, because a deleted row comes back the next time the scraper runs.
The transferable part
Most directory businesses treat content as something you buy or something you wait for users to contribute. There is a third option that is neither, and it scales: the businesses in your directory have already published, in a machine-readable format, the exact assets you need, specifically so that other people’s software can display them well.
That is not scraping in the adversarial sense. og:image is an invitation. The work is not extraction; it is the unglamorous part — the quality gate, the logo-versus-photo classifier, the deduplication that knows which duplicates are load-bearing, and the discipline to verify against the edge instead of the origin.
Nānā i ke kumu — look to the source. The data was already there, published by the people who know the business best. We just had to go read it, and then be honest about what we found.
Davis-Bacon vs. Hawaiʻi Chapter 104: the overtime math that creates back-wage findings
Hawaiʻi's prevailing wage law and the federal Davis-Bacon Act compute overtime differently — different triggers, and a different base. A sub who learned one and applies it to the other either underpays or overbids. Here is the actual math, worked on real wage-determination numbers, plus how to tell which regime governs your job.
Five products, one maritime spine
Binnacle started as crew credential tracking and turned into five products — fleet compliance, a pilot association dispatch board, harbor management, USCG exam prep, and an ARPA radar simulator. Why they share a data model, and why the regulated middle of an industry is a better place to build than the consumer edge.
What is a TMK number in Hawaii, and how do you look one up?
TMK stands for Tax Map Key — the unique parcel identifier used by all four Hawaiian counties. Every real estate transaction, permit application, and zoning lookup in Hawaii starts with the TMK. Here's how the system works and the fastest way to find one.