Stop asking which adjectives describe the brand. Show people the same fact written five ways and ask which one sounds like you. Voice guidelines built from choices, not vibes.
Every voice and tone project starts the same way. Somebody circulates a list of adjective pairs. Are we formal or casual, playful or serious, bold or measured. The org picks the flattering end of every axis. The output is a document saying the brand is confident, human, clear, and expert.
Nobody disagrees with it, and nobody can use it. Two writers reading "confident but human" produce completely different sentences, and when one of them gets edited, the argument that follows is about taste with no way to resolve it.
The deeper failure is that people can't reliably describe how they want to sound. They can, however, recognize it instantly when they see it.
The words in advertising should be like windows in a store.
David Ogilvy · via Eddie Shleyner, Very Good CopyTake one real fact the company actually has to communicate. Write it five ways across a boldness range. Ask people to pick the one that sounds most like the company. Repeat across a dozen scenarios. What comes back is a calibrated position, expressed in sentences you can hand a writer.
| Adjective survey | Calibration | |
|---|---|---|
| What you ask | "How bold should we be, 1–5?" | "Which of these five sentences would you sign?" |
| What people do | Pick the aspirational answer | Flinch at the one that goes too far |
| What you get | An average that describes nobody | A ceiling with a real sentence attached to it |
| What a writer does with it | Guesses | Matches the register of an approved example |
| What happens in review | Argument about taste | "That's a tier above what we agreed to" |
The last row is the one that changes daily life. When boldness is a documented tier with example sentences, a review comment stops being a personal preference and becomes a reference to a decision the group already made.
Pull actual claims the company has to make: a research finding, a product capability, a competitive position, a piece of bad news. Invented examples let people answer about a hypothetical brand. Real ones force them to answer about this one.
Four or five versions per scenario, spanning different positions rather than word-level variation. The range has to include one that's too safe and one that's too far, or the exercise only measures the middle.
Using a manufactured brand, Meridian, a field service operations platform, with the same underlying fact:
| Tier | The same fact, five ways |
|---|---|
| Measured | Teams reviewing work in a shared workflow reject more line items than teams reviewing after the fact. |
| Recommended | When the record shows up after the decision, you're not approving work. You're filing it. |
| Bold | Most field service teams don't have an approval process. They have a receipt collection process. |
| Boldest | You are not in control of your spend. You're just finding out about it politely. |
| Too far | If you're still approving repairs from a text message, you're the reason the budget is gone. |
The "too far" option is load-bearing. It's how you find the ceiling, and including it signals that the exercise is asking rather than steering toward a predetermined answer.
Voice is easy on a landing page and hard everywhere else. The calibration has to include the hard cases, because those are where writers guess:
Radio groups and dropdowns, keyboard-navigable, no timing, progress visible, answers copyable at the end. Response rate is the whole ballgame. A calibration with six responses is one person's taste with extra steps.
The output isn't an average. Averaging tiers produces the same mush the adjective survey did.
House voice has a default tier, but individual surfaces have ceilings. A crisis response capped at Measured stays there even if the campaign is running at Bold. Writing the cap table is the part that survives contact with a real quarter.
The guidelines that come out of this are structurally different from the adjective kind, because every rule carries an example that a group of humans already approved.
The instrument produces a document, but the exercise produces something more useful: shared authorship. Twenty people picked these sentences. When a draft gets pushed back for being off-voice, the writer is being pointed at a decision the whole team made.
That's the difference between guidelines that live in a folder and guidelines that live in review comments.
Writing five distinct versions of a dozen facts takes a day or two, and it has to be done by someone who can actually write at all five tiers convincingly. A weak "boldest" option biases the whole instrument toward the middle.
It's also worth being honest about the limit: this calibrates register and boldness. It doesn't decide what to say. That comes from the messaging architecture, and a calibration run before the conviction exists just produces confident-sounding sentences about nothing.
AI-Enabled Creative Operations covers how a calibrated voice gets encoded into a skill the whole team runs, so consistency stops depending on who happens to be drafting.