← Files DesignlyARCHIVED FILE
skills/creative-director/references/scoring-calibration.md
16.7 KB · Oct 4, 2026 · 12:30 UTC
# Scoring Calibration: Reference Rubric
This document is a calibration table for evaluating creative ideas. Three parallel scoring systems, calibrated against real campaigns. Use it as a reference during every evaluation to prevent score inflation and ensure reproducibility.
---
## 1. Six Criteria: Detailed Rubric
### Originality (weight 0.25)
The most significant criterion. Evaluates how novel the idea is for the category and culture at large.
- **1-2**: Pure category template. Any brand could have done this. Seen it 100 times. A bank with a handshake, a car on a mountain road, a family at the table with the product.
- **3**: Minimal attempt to differentiate, but it doesn't go beyond genre conventions. Different font, same thought.
- **4-5**: Noticeable attempt to stand out, but within expectations. Expected twist — the viewer guesses the payoff. "Not bad" for the category, but doesn't move culture.
- **6-7**: Fresh approach. Visible authorial thinking. The idea doesn't fit the category template. A competitor couldn't slap their logo on it without losing meaning.
- **8-9**: Unexpected. Nothing like this existed in the category. Remembered a week later. People retell it to each other. Burger King "Whopper Detour" — using McDonald's geolocation to sell their own burger.
- **10**: Creates a new genre, format, or approach. Changes the rules for the entire category. Dove Real Beauty redefined how brands talk about beauty. Old Spice "The Man Your Man Could Smell Like" created a language that was copied for years.
### Strategic Fit (weight 0.20)
Evaluates how well the idea solves the brief's objective and resonates with the target audience.
- **1-2**: Doesn't solve the brief's objective. The target audience won't recognize themselves. You could swap in any brand — nothing changes. The idea exists in a vacuum.
- **3**: Formally touches the brief's topic, but misses the target audience or the objective. Like answering the wrong question.
- **4-5**: Formally on-brief, but imprecise. Either the target audience is too broad/narrow, or the objective is addressed indirectly, or the insight is superficial.
- **6-7**: Clearly on-brief. Target audience is relevant, objective is addressed, insight works. Professional work that will satisfy the client.
- **8-9**: Perfectly on-brief AND adds something the client didn't ask for but needed. Expands the understanding of the task. Always #LikeAGirl — the brief was about the product, the solution became a cultural statement.
- **10**: Redefines the task. The client realizes they were asking for the wrong thing. Apple "Think Different" — not a computer ad, but a manifesto for those who think differently. Strategy above the brief.
### Emotional Response (weight 0.20)
Evaluates the strength and specificity of the emotional reaction.
- **1-2**: No reaction. A wall of text. Boredom. The viewer scrolls past without finishing.
- **3**: Minimal contact — the person noticed the ad but felt nothing. Background noise.
- **4-5**: Slight interest or a smile, forgotten within a minute. There's an emotion, but it's unnameable — "well, it's positive."
- **6-7**: Evokes a specific emotion. You can name it: curiosity, nostalgia, pride, surprise. The person remembers it for a day.
- **8-9**: Strong reaction: laughter, tears, anger, surprise. Tier 3 emotion — not "positive" but "lump in throat," "goosebumps," "want to show my mom." Thai Life Insurance ads — grown men cry.
- **10**: A cultural moment. People discuss, share, remember for years. John Lewis Christmas becomes an event people look forward to. Advertising transcends advertising.
### Feasibility (weight 0.15)
Evaluates the realizability of the idea given the budget, timeline, and resources.
- **1-2**: Unrealizable with the given budget, timeline, or resources. Requires technologies that don't exist or a budget 10x larger than available.
- **3**: Technically possible, but requires so many compromises that the idea loses its essence.
- **4-5**: Theoretically possible, but risky and expensive. Many dependencies, many things can go wrong. Production will eat the entire budget.
- **6-7**: Realizable with reasonable effort. Clear production plan, understood risks, adequate budget.
- **8-9**: Elegantly realizable. A simple solution to a complex problem. Minimal dependencies, maximum control. "Share a Coke" — just print names on cans.
- **10**: Can be executed tomorrow with minimal resources for maximum impact. The idea is so simple to execute that the main question is "why hasn't anyone done this before?"
### Scalability (weight 0.10)
Evaluates the idea's potential for development across time, channels, and markets.
- **1-2**: One-time execution. No series, no development. Done and forgotten.
- **3**: Can be repeated 1-2 times, but with diminishing returns. Second time won't surprise.
- **4-5**: 2-3 variations, but limited potential. Works in one channel. Adaptation to other markets is problematic.
- **6-7**: A season-long series. Works across 3+ channels. Can be adapted to other markets with localization.
- **8-9**: A platform for years. Infinite series. International potential without losing meaning. Spotify Wrapped — new every year, anticipated every year.
- **10**: A cultural platform that outgrows advertising. Red Bull Stratos — not a campaign, but a world-scale event. Nike "Just Do It" — 30+ years, infinite iterations, cultural code.
### Simplicity (weight 0.10)
Evaluates how quickly and easily the idea is understood.
- **1-2**: Requires instructions to understand. Complex mechanics. Three steps, two conditions, one QR code.
- **3**: Understandable, but with effort. Need to think, reread, ask again.
- **4-5**: Understandable after explanation. The idea doesn't read on its own, but once explained — it's logical.
- **6-7**: Immediately understandable, but requires 30 seconds of attention. There's a nuance that unfolds.
- **8-9**: One sentence. One image. Instant. Can be retold in 5 seconds and the listener gets it.
- **10**: So simple it seems obvious. But no one did it before. Coca-Cola "Share a Coke" — names on cans. That's it. Genius.
---
## 2. HumanKind Scale (Leo Burnett) — Detailed Scale
Leo Burnett's proprietary scale evaluating an idea's impact on people and culture. Use as a second filter after the six criteria.
| Score | Level | Description | Calibration Example |
|-------|-------|-------------|---------------------|
| 1 | Destructive | Actively harms the brand. People reject, boycott. Pollutes the media space. Tone-deaf. | Pepsi Kendall Jenner (trivializing protests) |
| 2 | No Idea | No thought. Formal budget fulfillment. Product in frame, logo at the end. | Typical stock-photo banner "20% off" |
| 3 | Invisible | No interest, no emotions. Wallpaper. Cliché. Viewer doesn't notice, doesn't remember, doesn't react. | Bank ad "we're close by" with a handshake |
| 4 | No Brand Purpose | There's some kind of idea, but it's unclear why the brand exists. Product without meaning, message without position. | Ad demonstrating the product but not answering "why" |
| 5 | Brand Purpose | There's a mission, people understand what the brand does and why. Meaningful communication, but no breakthrough. | IKEA "The Wonderful Everyday" |
| 6 | Intelligent Idea | Smart approach. Engaging. Not tied to a single channel. The idea is bigger than the format. | Burger King "Whopper Detour" |
| 7 | HumanKind Act | Changes people's thoughts, feelings, or actions. Flawless craft. Idea and execution at the highest level. | Nike "Dream Crazy" (Kaepernick) |
| 8 | Changes Thinking | Useful, interesting, becomes part of people's lives. People are grateful to the brand for the experience. | Spotify Wrapped |
| 9 | Changes Living | Inspires lifestyle change. People act differently because of the brand. | Dove "Real Beauty" (the entire campaign) |
| 10 | Changes the World | Social change on a global scale. Theoretical ceiling. | Theoretical ceiling — practically unachievable in advertising |
Important: most professional work falls in the 4-6 range. A score of 7+ is already Cannes shortlist level. Don't inflate.
---
## 3. Grey Scale — Grey Group Evaluation Scale
An alternative scale focusing on the quality of creative work as such.
| Score | Level | Description |
|-------|-------|-------------|
| 1 | Toxic | Harms the brand and people. Offensive, incompetent, dangerous. |
| 2 | Careless | Sloppy, unprofessional. Visible errors, no attention to detail. |
| 3 | Dull | Boring, predictable. Evokes no response whatsoever. |
| 4 | Expected | Expected for the category. "Well, okay." Not embarrassing to show, but nothing to be proud of. |
| 5 | Capable | Professional but unremarkable. Solid work without ambition. |
| 6 | Gratifying | Pleasant, you want to watch. There's something that hooks. Above average. |
| 7 | Original | Fresh, memorable. You want to show colleagues. Potential for local festivals. |
| 8 | Best in category | Best in category for the year. Competitors are envious. Cannes shortlist. |
| 9 | Best in show | Cannes shortlist / D&AD Pencil. Work the industry talks about. |
| 10 | Best in the world | Grand Prix. Sets the standard for years to come. |
The median for a quality agency is 5-6. If your scores are consistently higher — you're not calibrating, you're flattering.
---
## 4. Calibration Anchors: Real Campaigns
These campaigns are fixed points on the scale. Use them for comparison during every evaluation.
**Level 10 (Grand Prix / cultural shift)**
- Fearless Girl (McCann, State Street) — sculpture facing the bull on Wall Street. Simplicity, symbolism, cultural resonance. Became a permanent monument.
- Dumb Ways to Die (McCann Melbourne, Metro Trains) — a song about railway safety. 5 billion views, 21% reduction in incidents. Advertising that became pop culture.
**Level 9-9.5 (Cannes Gold / Lion / D&AD Pencil)**
- Dove Real Beauty Sketches (Ogilvy) — FBI sketch artist draws women from their own description vs. a stranger's description. Visualizing distorted self-perception.
- Burger King Moldy Whopper (INGO, David, Publicis) — a beautiful shot of a moldy burger. Shock + message about no preservatives.
- Nike Dream Crazy (Wieden+Kennedy) — Kaepernick: "Believe in something, even if it means sacrificing everything." Polarization as strategy.
**Level 8-8.5 (Cannes Silver-Bronze / best in category)**
- Spotify Wrapped (in-house) — personalized year-end summaries. Users create content themselves. An annual tradition.
- IKEA ThisAbles (McCann Tel Aviv) — 3D-printed furniture adapters for people with disabilities. Free, open-source.
- Always #LikeAGirl (Leo Burnett) — reframing the insult "like a girl" into a compliment.
**Level 7-7.5 (solid professional work, local awards)**
- Good work from a regional agency. Clear insight, professional execution, but no cultural breakthrough. Local festival shortlist.
**Level 6 and below (category average)**
- Category standard. Professional but predictable. Not remembered after a week.
---
## 5. Rules Against Score Inflation
Inflation is the main enemy of useful evaluation. AI models tend to inflate scores. Apply these rules strictly.
**Rule 1: Real Analogues Test.** Before assigning a score of 9+, name 3 real Cannes-winning campaigns at that level. Is your idea truly comparable? If not — lower the score.
**Rule 2: Batch Control.** If all ideas in a batch score 8+, something is wrong with calibration. This is statistically impossible. Re-evaluate the entire batch from scratch.
**Rule 3: Normal Distribution.** The average score for a batch of 10 ideas should be 5-6. If the average is above 7, you're inflating. Go back and recalibrate.
**Rule 4: Specificity Test.** Replace the brand with a competitor. If the idea works exactly the same — originality is 5 at most. The idea must be inseparable from the brand.
**Rule 5: "Would I share this?" Test.** Mentally show the idea to a friend who doesn't work in advertising. If they're not impressed, won't forward it, won't mention it at dinner — score below 7. Ad people tend to overvalue "clever" ideas that only work within the industry.
**Rule 6: First Reaction.** Record your initial score before analysis. If the score rises more than 2 points after analysis — you're rationalizing, not evaluating. Return to your first reaction.
**Rule 7: Time Test.** Imagine this idea a year from now. Will it be in a "best of the year" compilation? If not — no higher than 7.
---
## 6. Multi-Perspective Panel (CAT Simulation)
Consensual Assessment Technique (CAT) — a method where multiple experts evaluate independently, then compare results. Simulate four perspectives.
### Four Perspectives
**Creative Director (craft and idea).** Focus: originality of approach, execution quality, visual language, tone of voice. Question: "Would I be proud of this work in my portfolio?"
**Strategist (brief fit and insight).** Focus: brief alignment, insight depth, target audience fit, business logic. Question: "Does this solve the client's problem?"
**Consumer (interest and emotion).** Focus: first reaction, comprehensibility, emotional response, desire to engage. Question: "Is this interesting to me? Would I stop scrolling?"
**Jury Member (award-worthiness).** Focus: novelty for the industry, craft, scale of the idea, case potential. Question: "Will this make the Cannes shortlist?"
### Panel Rules
1. Each perspective evaluates independently, without seeing others' scores.
2. The final score is not an arithmetic average, but the result of discussing discrepancies.
3. If the discrepancy between perspectives is greater than 2 points — it's a signal. Investigate the cause.
4. High discrepancy = polarizing idea. This can be good for awards (juries love boldness), but risky for the client (consumers may not accept it).
5. Low discrepancy with high scores = strong idea with consensus. Rare.
6. Low discrepancy with average scores = safe but unambitious work.
### Typical Discrepancy Patterns
- CD high, consumer low: an "advertising" idea that the industry loves but people don't.
- Consumer high, jury low: a mass idea without creative novelty. It works but doesn't surprise.
- Strategist high, CD low: strategically sound but dull in execution. Needs better craft.
- Jury high, strategist low: "festival" work created for awards, not for business.
---
## 7. Three Axes: Final Summary
The final evaluation is built on three axes. Each axis serves its own function.
### Axis 1: Brief Compliance (gate)
This is a binary filter, not a scale. The idea either passes or doesn't. Any failure on this axis means the idea is not admitted to further evaluation, regardless of creative strength.
Checkpoints: target audience matches the brief, objective is addressed, brand is recognizable, constraints are met, tone of voice is appropriate.
### Axis 2: Idea Strength (6-criteria weighted score)
Weighted score across six criteria:
```
Total = Originality * 0.25
+ Strategic Fit * 0.20
+ Emotional Response * 0.20
+ Feasibility * 0.15
+ Scalability * 0.10
+ Simplicity * 0.10
```
Interpretation: 8+ excellent, 6-7 good, 5-6 average, below 5 weak.
### Axis 3: Scalability and Idea Level
Four questions to assess scale:
1. How many channels does the idea cover without adaptation?
2. How many iterations/series can be produced?
3. Does the idea work in other markets?
4. Can the idea outlive a single flight?
### Idea Level Matching
It is critically important to determine the idea's level and match it to the brief's objective.
**Big Idea** — a platform idea that defines the brand for years. Example: Dove "Real Beauty," Coca-Cola "Open Happiness," Nike "Just Do It."
**Campaign Idea** — a specific campaign within a Big Idea. Example: Dove "Real Beauty Sketches" (within the Real Beauty platform), "Share a Coke" (within Open Happiness).
**Execution Idea** — a specific execution within a campaign. Example: Dove "Girls Unstoppable" (within Sketches), Coca-Cola "Unlock the 007 in You" (within Share a Coke).
### Mismatch as a Problem
The idea's level must match the brief's objective. Level mismatch is a common error:
- Big Idea for shelf talkers — a waste of time and resources. The client needs an execution, and they got a manifesto.
- Execution Idea for a rebrand — too small. The client needs a platform, and they got a single poster.
- Campaign Idea when a Big Idea is needed — will work for six months and die. No foundation for growth.
Before evaluating, determine: what idea level does the brief call for, and what idea level is being proposed. If there's a mismatch — it lowers the strategic fit score, even if the idea itself is strong.
---
*Author: Serge Shima ([t.me/aimastersme](https://t.me/aimastersme) · [sergeshima.com](https://sergeshima.com) · [aimasters.me](https://aimasters.me)) · License: CC BY 4.0 — attribution required · Source: [smixs/creative-director-skill](https://github.com/smixs/creative-director-skill)*
SHA-256: 3b92369582c2618f3385b8ffcdf22223ad136cf1d7b1c91aba4b0942f17708c8