ZachSearcy
← All frameworks
Framework How to actually build voice guidelines

Voice Calibration

Stop asking which adjectives describe the brand. Show people the same fact written five ways and ask which one sounds like you. Voice guidelines built from choices, not vibes.

The problem it solves

Every voice and tone project starts the same way. Somebody circulates a list of adjective pairs. Are we formal or casual, playful or serious, bold or measured. The org picks the flattering end of every axis. The output is a document saying the brand is confident, human, clear, and expert.

Nobody disagrees with it, and nobody can use it. Two writers reading "confident but human" produce completely different sentences, and when one of them gets edited, the argument that follows is about taste with no way to resolve it.

The deeper failure is that people can't reliably describe how they want to sound. They can, however, recognize it instantly when they see it.

The words in advertising should be like windows in a store.

David Ogilvy · via Eddie Shleyner, Very Good Copy
You look through them. Ogilvy's point was that a headline someone compliments is usually a headline that failed, because they noticed the writing instead of the offer. It's the argument for why voice work has to be calibrated against real sentences carrying real claims. Adjective surveys optimize for how the brand wants to be described. Choosing between actual sentences optimizes for whether the claim gets through.
The core move

Take one real fact the company actually has to communicate. Write it five ways across a boldness range. Ask people to pick the one that sounds most like the company. Repeat across a dozen scenarios. What comes back is a calibrated position, expressed in sentences you can hand a writer.

Why choices beat adjectives

Adjective surveyCalibration
What you ask"How bold should we be, 1–5?""Which of these five sentences would you sign?"
What people doPick the aspirational answerFlinch at the one that goes too far
What you getAn average that describes nobodyA ceiling with a real sentence attached to it
What a writer does with itGuessesMatches the register of an approved example
What happens in reviewArgument about taste"That's a tier above what we agreed to"

The last row is the one that changes daily life. When boldness is a documented tier with example sentences, a review comment stops being a personal preference and becomes a reference to a decision the group already made.

Building the instrument

1. Use real facts, never invented ones

Pull actual claims the company has to make: a research finding, a product capability, a competitive position, a piece of bad news. Invented examples let people answer about a hypothetical brand. Real ones force them to answer about this one.

2. Write the same fact across the full range

Four or five versions per scenario, spanning different positions rather than word-level variation. The range has to include one that's too safe and one that's too far, or the exercise only measures the middle.

Using a manufactured brand, Meridian, a field service operations platform, with the same underlying fact:

TierThe same fact, five ways
MeasuredTeams reviewing work in a shared workflow reject more line items than teams reviewing after the fact.
RecommendedWhen the record shows up after the decision, you're not approving work. You're filing it.
BoldMost field service teams don't have an approval process. They have a receipt collection process.
BoldestYou are not in control of your spend. You're just finding out about it politely.
Too farIf you're still approving repairs from a text message, you're the reason the budget is gone.

The "too far" option is load-bearing. It's how you find the ceiling, and including it signals that the exercise is asking rather than steering toward a predetermined answer.

3. Cover the scenarios where voice actually breaks

Voice is easy on a landing page and hard everywhere else. The calibration has to include the hard cases, because those are where writers guess:

  • Delivering a research finding that implicates the reader
  • Naming a competitor's weakness without naming the competitor
  • Admitting a limitation of your own product
  • Responding to criticism from the community you serve
  • Announcing good news without a victory lap
  • Writing to a technical buyer and an executive buyer about the same thing

4. Make it take ten minutes and require no mouse

Radio groups and dropdowns, keyboard-navigable, no timing, progress visible, answers copyable at the end. Response rate is the whole ballgame. A calibration with six responses is one person's taste with extra steps.

Reading the results

The output isn't an average. Averaging tiers produces the same mush the adjective survey did.

  • The mode is the default tier. Where most people land is where house voice sits.
  • The highest tier anyone picked without an outlier cluster is the ceiling. The ceiling, not the average. That's what "how bold can we go" actually means.
  • Disagreement by role is a finding, not noise. If sales picks two tiers hotter than legal every time, that's a real constraint to design around, and it's better to know it now than in a review cycle.
  • Scenario-level variance sets the cap table. People are consistently bolder on research findings than on crisis response. That variance becomes surface-specific caps.
Caps always win

House voice has a default tier, but individual surfaces have ceilings. A crisis response capped at Measured stays there even if the campaign is running at Bold. Writing the cap table is the part that survives contact with a real quarter.

What you ship

The guidelines that come out of this are structurally different from the adjective kind, because every rule carries an example that a group of humans already approved.

  • A tier ladder, four named boldness levels, each with three to five real sentences from the calibration
  • A cap table, which surfaces are allowed which tiers, and what always wins in a conflict
  • Approved openings, the actual first lines people chose, by format
  • Banned patterns with replacements, so a writer knows what to use instead of a forbidden word
  • A review rubric, so "does this sound like us" gets scored rather than argued

The organizational effect

The instrument produces a document, but the exercise produces something more useful: shared authorship. Twenty people picked these sentences. When a draft gets pushed back for being off-voice, the writer is being pointed at a decision the whole team made.

That's the difference between guidelines that live in a folder and guidelines that live in review comments.

What it costs

Writing five distinct versions of a dozen facts takes a day or two, and it has to be done by someone who can actually write at all five tiers convincingly. A weak "boldest" option biases the whole instrument toward the middle.

It's also worth being honest about the limit: this calibrates register and boldness. It doesn't decide what to say. That comes from the messaging architecture, and a calibration run before the conviction exists just produces confident-sounding sentences about nothing.

Next

Guidelines only hold if something enforces them at the moment of writing.

AI-Enabled Creative Operations covers how a calibrated voice gets encoded into a skill the whole team runs, so consistency stops depending on who happens to be drafting.