Cognitive Boundary Breaker · AI that does more than agree

See one more direction before an important decisionv2.0 · Move quickly on small things; think clearly on big ones

It surfaces hidden assumptions, checks the evidence, and puts the strongest objection on the table. It does not decide for you; it helps you avoid one more untested leap.

Not every question needs a thinking marathon. Routine questions should stay quick. Important, hard-to-reverse choices deserve one more careful pass.

Platform facts verified as of 2026-08-27 Free to share · Please keep the source when you can GO · Reinventing Organizations

Start here

Try it once before you install anything

No learning curve and no code. Copy one short prompt into the AI tool you already use.

30-second trial

Send this with your question

Useful before a meeting, a proposal review, a career choice, or any moment when agreement is not enough.

30-second trial
I have an important question. Do not agree with me first or jump straight to an answer. Please: 1) identify no more than three key assumptions; 2) give the strongest counterargument; 3) separate facts, inferences, and guesses; and 4) propose one minimum validation step.

My question is:
Keep using it

Use it in your AI tools

No coding needed—set it up in ChatGPT, Claude, DeepSeek, or Kimi.

Choose my AI

Blind comparison

Same question, same model, a different reasoning scaffold

These five fictional cases show how one model's reasoning path can change after reading the Skill. They are single-output examples—not an average uplift and not evidence that the facts are more accurate.

Single simulated outputs · same default model · no web or tools · ordinary answers were not weakenedView the old demonstration's raw responses

Original question

Stay or join the startup?

I am 35 and work in management at an established company. My income is stable, but I have not grown for two years. A startup has invited me to join for 25% less cash salary, promised equity with terms still undefined, and a heavier workload. I have no mortgage and twelve months of savings. Should I go?

Ordinary responseNo Cognitive Boundary · single pass

Do not decide yet. Put fully diluted ownership, vesting, exercise and dilution terms in writing; verify runway, paying customers, authority and six-month success metrics. Go only if clear entry conditions hold and the role is worthwhile even if the equity becomes worthless.

With Cognitive BoundarySame default-model environment · full method

Value the equity at zero until terms are written. Ask whether 25% less cash buys rare skills, real scope and credible upside. Go only if you can absorb 18 months, retain nine months of reserves, and leave with portable results; otherwise run a 90-day search.

What changed
Equity at zeroDownside testFallback plan

These are fictional comparison cases. Health content is not medical advice; investment content is not personalised financial advice.

Prompts and excerpts are faithful adaptations of the Chinese originals. Excerpts are concise summaries; the complete raw responses remain in Chinese.

Re-evaluation in progress

We will not declare victory for ourselves

In the first five-case sample, ordinary answers were already strong and the gap was too small to support a claim of significant improvement. We removed the homepage scores and replaced them with a stricter three-condition evaluation.

Current evidence boundaryEfficacy conclusion: not establishedNew protocol in progress

Problem space

Does the answer test the initial frame and expose the real decision, constraints and assumptions?

Evidence space

Does it separate what is known, inferred and missing, then identify useful checks?

Possibility space

Does it reveal genuine options, staged paths, decision rules and ways back?

Counterfactual space

Does it state the strongest opposing case, failure scenario and evidence that would change the conclusion?

The new plan covers 12 simulated cases, three response conditions, 56 pre-specified checks, and three independent blinded raters. Failures, ties, negative controls, and response cost will remain visible. No average uplift, win-rate, or safety claim will be published before completion.

Open the new protocol and old sample

The new evaluation compares an ordinary answer, a three-line lightweight prompt, and the full Skill across engineering, life, health, investing, learning, and low-stakes questions. It prioritises critical omissions, counterevidence, stopping rules, safety errors, and response cost instead of a single attractive composite score.

Mechanism, not magic

Why a reasoning scaffold can change the answer

A language model has no fixed internal decision procedure. What it checks depends heavily on the context supplied before the answer settles.

What usually happens

  1. The model predicts a plausible next token from the preceding context.
  2. It often adopts the user's framing, tone and implied goal.
  3. It develops a fluent answer that appears useful and coherent.
  4. It stops when the response sounds sufficiently complete.
The starting frame may remain untested.Agreement can feel safer than correction.Evidence and inference can sound equally certain.No falsifier, threshold or next test is automatically required.

What the Skill adds

  1. Triage stakes, reversibility and the decision window.
  2. Reframe the problem and label assumptions and evidence.
  3. Construct the strongest opposing case.
  4. Set thresholds, a falsifier and the smallest useful test.
  5. End with action, owner, timing, rollback and a learning record.

The Skill does not make the model know more. It directs attention toward checks that ordinary prompting often leaves implicit.

A loop, not an assembly line

Eight stations form a loop you can enter from any point. Only three rules hold the structure together: triage sets the depth; experiments turn thought into evidence; and the persistence layer carries the learning into the next round.

Question Triage Stakes × Reversibility × Window Low stakes / Reversible Medium stakes High stakes · Irreversible Fast Track Best answer + top three risks Stop here Standard Track Default: Question the Question + Active Falsification Deep Track · Loop (enter anywhere) DELIBERATION ZONE 1 Question the Question 2 First-Principles Decomposition 5 Paradigm Reconstruction EVIDENCE ZONE 3 Global Scan 4 Future Backcasting 6 Active Falsification Only Exit Station 7 · Experiment 7/30/90 days · Lock metrics in advance Result + deadline Station 8 · Persistence Layer Assumption ledger · Decision log · Calibration The first action of the next cycle = review the previous cycle's ledger, not ask a new question
The red loop is what separates this system from one unusually good conversation. Remove it and the eight stations become a long piece of analysis. Connect it and assumptions can be closed, confidence can be calibrated, and the cycle can get shorter. If a loop never produces a “closed assumption,” it is not a learning system. It is a writing exercise.
01

Next-token prediction

A language model generates text token by token from probabilities conditioned on context. A fluent continuation is not automatically an audit of the decision.

02

Alignment and sycophancy

RLHF rewards helpful, safe and instruction-following replies. It can also reinforce agreement with a user's framing or preference. That is a tendency, not an inevitability.

03

Prompt-time scaffold

The Skill places a checking procedure in context. It directs attention toward assumptions, evidence gaps, countercases and tests; it does not retrain the model or add facts.

What this does not prove

  • A scaffold cannot supply missing facts. Without tools, external claims remain unverified.
  • Five cases, one model, one date and one pass do not establish universal or causal improvement.
  • More structure can overanalyse simple, reversible choices. It also does not replace timely clinical assessment or independent professional investment advice.
Sources and further reading

The references below support the general mechanisms of language models, instruction following, RLHF and sycophancy—not the particular score of this small benchmark.

  1. Attention Is All You Need
  2. Language Models are Few-Shot Learners
  3. Training language models to follow instructions with human feedback
  4. Towards Understanding Sycophancy in Language Models
  5. TruthfulQA: Measuring How Models Mimic Human Falsehoods
  6. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
  7. Self-Consistency Improves Chain of Thought Reasoning in Language Models

How it works

Three simple habits for better judgment

Rigor does not have to feel academic. The system turns three commonly missed steps into repeatable actions.

01

Ask the right question

Surface hidden assumptions and make sure you are solving the problem, not just its symptom.

02

Look for uncomfortable evidence

Separate facts from inference and actively seek counterexamples, not only supporting material.

03

Return to the real world

Turn the conclusion into a small test, a deadline, and an observable result.

Open the full method: eight-station loop, evidence levels, and prompt libraryThe complete method, countercase protocol, hypothesis ledger, and evidence register remain here when you need the depth.
01

v1.0 was a method. v2.0 is a system.

v1.0 helped you ask better questions. In practice, three other things still get in the way—and none can be fixed by polishing the prompt alone.

Not every question needs a deep dive

v1.0 sent every question through all six layers. In practice, people either stopped using it or performed the ritual without gaining much.

v2.0 adds a triage layer: the stakes × reversibility decide how deep to go. It also adds an anti-paralysis rule—if analyzing a reversible decision will take longer than doing it and rolling it back, stop analyzing and act.

What AI says ≠ what exists in the world

v1.0 asked for a “global scan.” But without real tools, a model can only produce something that looks comprehensive.

v2.0 adds L1/L2/L3 evidence grades and a citation spot-check protocol. Conclusions may rest only on L2 or L1 material. If the model cannot search, the whole station steps down and says so at the top of the answer.

A breakthrough today ≠ a lesson remembered next time

v1.0 lived inside a single conversation. There was nowhere to record what changed, so the learning loop never closed.

v2.0 adds a persistence layer: an assumption ledger, a decision log, and a calibration record. If assumptions are not written down and predictions are not scored, “updating your thinking” remains a good intention.

04

A prompt library for working in rounds

Depth comes from multiple rounds, with a person choosing what matters between them. A one-shot prompt has only one valid job: quickly sketch the problem map. Let the triage result determine which station comes next. Human judgment is the router between rounds, and that role should not be handed to AI.

0TriageOutput = Track choice and rationale

Before any deep analysis, answer a simpler question: how much thought is this decision worth? Triage is not ceremony. It prevents two equal failures—running the full framework on everything until it becomes a ritual, and never using it at all.

Triage Prompt
Before beginning the analysis, please triage this issue.

Assess:
1. How high are the stakes of this decision? (Resource commitment, scope of impact, duration of impact)
2. Is this decision reversible? What would reversal cost?
3. How long is the decision window?
4. Is this a one-time issue, or a structural issue that will recur?

Then recommend:
- Fast track: directly give your best answer and the top three risks;
- Standard track: recommend which 2–3 stations to activate and explain why;
- Deep track: recommend entering the complete loop and explain which station is the most effective entry point.

If there is not enough information to triage, ask me questions first. Do not assume.
1Question the questionOutput = one clearly reframed question

Before asking AI for an answer, ask whether the question deserves to be framed this way. Keep the boundary clear: a useful result is one better question, not an essay listing everything wrong with yours.

Question the Question + Convergence
Do not answer my question directly.

First analyze:
1. What hidden assumptions does this question contain?
2. Which assumptions may be wrong, outdated, or unverified?
3. Have I mistaken an appearance, symptom, or result for the real problem?
4. What might be the original problem behind this question?
5. If you redefined the problem, how would you state it?

Prioritize helping me ask the right question over rushing to provide an answer.

After the analysis, you must converge and provide:
1. The redefined question (one sentence);
2. Its key differences from the original question (no more than three);
3. If this redefinition is wrong, where it is most likely wrong.
2Break it down from first principlesOutput = a list of facts, constraints, and the real goal

Strip away industry habit, history, and convention until only the basic facts, constraints, and true objective remain. First-principles thinking often fails by tearing down a convention that quietly protects a real constraint. That is why every convention marked for removal must first pass a Chesterton’s fence check: understand why the fence was built before taking it down.

First-Principles Decomposition + Fence Check + Verifiability
Analyze this problem from first principles.

Requirements:
1. Do not accept "the industry has always done it this way" as a reason;
2. Do not assume that existing organizations, processes, roles, product forms, and business models must exist;
3. Continue decomposing the problem until it can no longer be broken down;
4. Distinguish basic facts, technical constraints, resource constraints, institutional constraints, and human conventions;
5. Identify the true objective function of the system;
6. Give a conclusion that reconstructs the problem from first principles.

【Chesterton's Fence Check】For every industry convention or existing practice you recommend removing, first answer:
1. Why was this "fence" erected in the first place? What did it protect?
2. Does the constraint it protected still exist today?
3. Classify the convention accordingly: protective convention (the constraint remains real; do not dismantle) / obsolete convention (the constraint once existed but has disappeared; may dismantle) / inertial convention (never corresponded to a real constraint; may dismantle).
Only the latter two may enter a "remove" recommendation. If you cannot answer Question 1, treat the convention as protective by default.

【Verifiability of Constraints】For every "basic fact" and "real constraint" you list, label:
- Is it physically/mathematically necessary, or only temporarily true under current technology and cost conditions?
- How can it be verified, and how much would verification cost?
3Scan the wider worldOutput = a solution map labeled L1/L2/L3

This station depends on tools, not clever wording. No prompt can create genuinely exhaustive research. Without real search tools, a model draws “global examples” from what it remembers from training; outdated, distorted, or invented material can slip in and look every bit as credible as a real case.

Global Scan (with Evidence-Grading Requirements)
Use real retrieval tools to systematically scan solutions related to this problem around the world.

The search scope includes: academic papers and reviews, university laboratories, startups and new products, established-company cases, open-source projects, patents, industry reports, policy pilots, failure cases and controversies, and cases from the non-English-speaking world.

Complete the following:
1. Classify existing solutions;
2. Explain what problem each solves, its conditions for viability, costs, and limitations;
3. Identify gaps that remain unsolved;
4. Extract transferable underlying mechanisms;
5. Distinguish mainstream consensus, minority views, and frontier exploration.

Look especially for:
- Cases that run counter to mainstream practice but work;
- Mechanisms already solved in other industries but not yet adopted in this one;
- Solutions that failed before but may become viable again because technological conditions have changed;
- Small-scale experiments that work but have not yet scaled;
- Overlooked cases from the non-English-speaking world.

Evidence requirements:
- Attach a verifiable source to every critical case and data point;
- Clearly label which content comes from actual retrieval (L2) and which comes from your prior knowledge (L3);
- List any content that cannot be verified through retrieval separately as a "lead to verify"; do not mix it into conclusions.

If you do not have real retrieval capability in this conversation, say so directly and label all output from this cycle as L3 leads.
4Look back from the futureOutput = trend judgments with leading indicators and signals that could prove them wrong

Many problems exist only under today’s technology, costs, and rules. Change a key constraint and the original problem may disappear. But future thinking easily turns into a grand story that can never be wrong—and is therefore never useful. Every trend judgment needs an early indicator and a signal pointing the other way.

Future Backcasting + Falsifiability
Look back at today from the future.

Assume the points in time are 3, 5, and 10 years from now, and analyze:
1. Will today's problem still exist?
2. Which current constraints will disappear because of technological progress?
3. Which capabilities will become inexpensive, standardized, or nearly free?
4. What new scarce resources will emerge?
5. Which decisions seem reasonable today but will become path dependencies in the future?
6. If we backcast from the future, what capabilities should we build in advance today?

For every trend judgment you make, add:
1. What observable leading indicator should appear within the next 6–12 months?
2. What signal would show that this trend judgment is wrong?
3. Is this judgment a structural change / cyclical fluctuation / short-term noise? On what basis?

Do not give any trend statement that no evidence could overturn.
5Redesign the paradigmOutput = ≥3 system designs that change genuinely different dimensions

A real breakthrough often changes the system until the old problem no longer exists. Asked for 3 different paradigms, a model may offer three versions of the same structure. Add a diversity rule: turning the same lever to three different settings still counts as one idea.

Paradigm Reconstruction + Enforced Diversity
Do not optimize the current solution.

Try to overturn the old paradigm on which the current problem depends, and answer:
1. Could a new system make this problem cease to exist?
2. If you did not preserve the current organization, processes, roles, and business model, how would you design from scratch?
3. Which elements of the current system are merely historical legacies rather than necessary conditions?
4. How would a new company with no legacy burden solve it?

Propose at least 3 new paradigms subject to a diversity constraint: each paradigm's core dimension of change must differ, selecting one each from—mode of value creation / mode of delivery / transaction and pricing structure / allocation of responsibility and risk / time structure (one-time→continuous; synchronous→asynchronous).

If two paradigms are merely different intensities of the same lever, count them as one and regenerate.

Each paradigm must include:
1. Conditions for viability;
2. One minimum-cost validation experiment;
3. The strongest argument against it.
6Try to prove it wrongOutput = Tier 2 and Tier 3 opposing evidence

The Tier 1 prompt—an opposing view from another model session—is in the “Other Platforms” tab because it must run in a fresh conversation, preferably on a different model. This section covers Tiers 2 and 3. For a deep-track decision, Tier 1 alone does not count as completed falsification.

Tier 2 Opposition (Base Rates) + Tier 3 Opposition (Premortem)
【Tier 2 Opposition · Base Rates】

Build an outside view for this judgment:
1. What is the historical base success rate for decisions of this kind? (Attach retrieved sources)
2. Why should we believe we are the exception? Which of those reasons were also believed by everyone who failed?
3. Retrieve 3–5 failure cases most similar to ours and explain at which step each died.

【Tier 3 Opposition · Premortem】(This template should be completed with real team members. AI is only the facilitator and recorder, not the sole opponent.)

Assume this project has failed completely two years from now. Backcast in the form of a "failure review":
1. The 10 most likely causes of failure;
2. Which risks are already showing early signals today;
3. Which risks will the team deliberately ignore? Why is the team motivated to ignore them?
4. Which metrics look good but may actually represent false prosperity?
5. Which three assumptions most urgently need validation?
7Experiment and actOutput = experiments with deadlines and pass/fail measures

This is the only exit from reflection. Lock the decision measures before the result arrives, so the story cannot be rewritten afterward. Test the assumptions that could break the plan before optimizing its details. If an experiment reaches its deadline without a result, that absence is itself something to review.

Convert into Verifiable Action
Convert the discussion into:
1. The 3 most critical current judgments;
2. The critical assumption behind each judgment;
3. Validation experiments that can be completed within 7 days, 30 days, and 90 days;
4. Objective decision metrics for each experiment (locked in advance to prevent post hoc explanations);
5. Which actions are reversible and which are irreversible;
6. The next step with the lowest cost and greatest information gain.

Rules: every experiment must be entered in the assumption ledger with a deadline and decision metrics; experiments that falsify critical assumptions take priority over experiments that optimize solution details; an experiment that reaches its deadline without a result is itself a signal requiring review.
Know when to stopOutput = keep analyzing, or move into action

Stop the analysis as soon as any one condition is met: another round is unlikely to change the leading option; an experiment can resolve the remaining uncertainty more cheaply; or the decision window is closing and delay now costs more than further analysis is worth.

Convergence Judgment
Judge whether the current discussion should converge:
1. If we conduct one more round of analysis, what new information is most likely to emerge? Would it change the current preferred option?
2. For the largest current uncertainty, is continued analysis cheaper, or is running an experiment cheaper?
3. Given the decision window, do you recommend continuing analysis (analyzing what), or moving to action (doing what)?
ScoutOne-shot scouting promptOutput = a field map, all treated as L3

A master prompt that asks for everything at once conflicts with the rule against one-shot answers. v2.0 therefore limits it to scouting: use it to sketch the terrain quickly, never as the basis for a conclusion. Afterward, return to triage and work in rounds.

Scout Mode
【Scout Mode: Treat all content in this output as L3 leads. Use it only to map the problem, never as a conclusion.】

I want you not to answer within the boundaries of my existing cognition, but to help me see the full landscape of this problem.

Briefly complete the following in order (no more than 200 characters per section):
1. The problem's hidden assumptions and possible directions for redefinition;
2. Irreducible facts and constraints at the first-principles level;
3. Categories of solutions that may exist worldwide (label: not verified through retrieval);
4. Trends that may change the nature of the problem over the next 3–10 years;
5. At least 3 logically distinct new-paradigm directions;
6. The strongest opposing case's points of attack;
7. Which 2–3 stations you recommend I investigate first, and why.

This is not the final analysis; it is a campaign map.
05

The persistence layer: three simple tables

This is v2.0’s most important addition, and the one part no beautifully written prompt can replace. If assumptions are not recorded, predictions are not scored, and decisions are not revisited, “learning” is only a wish: every conversation begins at zero and the loop never closes. Keep the tables somewhere you will actually maintain—a document, a spreadsheet, or the platform’s built-in memory.

Three Persistence-Layer Tables + Loop-Closure Prompt (Markdown, ready to paste into a document)
## Assumption Ledger
| Assumption | Date proposed | Current confidence | Supporting evidence (grade) | Opposing evidence | Validation experiment | Deadline | Result |
|---|---|---|---|---|---|---|---|
|  |  |  |  |  |  |  |  |

Rule: add or update at least 3 entries in every deep cycle; every entry must have a deadline—an assumption without a deadline is a belief, not an assumption; do not delete falsified assumptions, mark them closed and retain them. They are the most expensive assets.

## Decision Log
| Decision | Date | Reasons at the time (≤3) | Expected result and metrics | Review date | Actual result | Cognitive update |
|---|---|---|---|---|---|---|
|  |  |  |  |  |  |  |

Rule: reasons must be recorded at the time of the decision—reasons reconstructed afterward always look wise; set the review date on the day of the decision and put it on the calendar.

## Calibration Record
| Prediction | Date | My confidence | Actual result | Difference |
|---|---|---|---|---|
|  |  |  |  |  |

Rule: review quarterly—among the things you said you were 80% confident about, what percentage actually happened? The gap between confidence and hit rate is how much your cognitive system needs to be corrected.

## Closing the Loop: Paste This at the Start of Every Cycle
Below are the assumption ledger and validation results from my previous cycle.
First complete: 1. Which assumptions have been confirmed / falsified / left untested past their deadlines? 2. Based on these results, which of my previous judgments should be revised? 3. How does the revised cognition change the issue to be discussed in this cycle?
Only after completing this review should you begin a new round of analysis.
How can you tell whether this system is working?Do not judge it by how impressive the documents sound. Track three measures: Calibration—is the gap between your confidence and your actual hit rate getting smaller? Assumption retirement rate—does each quarter end with some assumptions disproved and closed? If nothing is ever overturned, the falsification step is only theatre. Cycle time—is the average time from writing an assumption to testing and closing it getting shorter?
06

Where the method breaks

These are not footnotes. They are the specific ways the method can fail—moments that look like a breakthrough in thinking but are not.

Fluent language ≠ a correct conclusion

Require the model to separate known facts, evidence-based inferences, reasonable guesses, and assumptions with no evidence yet. The smoother the prose, the more carefully you should check it.

Found in search ≠ verified

For any citations that support the conclusion, spot-check the 3–5 most important. Open the original and ask: does it exist, was it represented accurately, and is it still current? If any check fails, downgrade it to L3 and search again.

A model’s objection ≠ completed falsification

The same model, with the same context and assumptions, tends to offer gentle objections. It creates the comfort of having been “challenged” without delivering a serious challenge. Tier 1 opposition is only a warm-up.

Use it for everything = stop using it

Maximum challenge on every question is as damaging as no challenge at all. A system told to “question everything” will manufacture doubt around sound, ordinary answers and add friction to simple choices. Doubt needs calibration too.

Do not let analysis replace action

A loop that produces documents but never closes an assumption is procrastination in formal dress. When any convergence criterion is met, act.

More models ≠ more perspectives

Mainstream models learn from heavily overlapping material. Five assigned roles may disagree less than you expect—and may agree inside the same blind spot. Independent information sources > model count. Once the disagreement is made visible, the final judgment belongs to a person.

07

Sources behind the platform claims on this page

This ledger is not here to look impressive. It tells you how far each sentence can be trusted. L1 means an official source or first-hand verification; L2 means traceable evidence that can be cross-checked; L3 is only a lead for the next search. All facts were checked on 2026-08-27. Platforms can change at any time, so read the grade, the date, and the interface in front of you together.

FactGradeSource
Claude skill package: the ZIP root must be the skill folder itself, with no additional wrapper folder;name ≤64 characters, description ≤200 charactersL1 Factsupport.claude.com
Claude Skills require code execution to be enabled first; available on Free / Pro / Max / Team / Enterprise; entry point: Customize → SkillsL1 Factsupport.claude.com
Developer-side Agent Skills specification: description ≤1024 characters (a different limit from the 200-character web upload; do not mix them up)L1 Factplatform.claude.com
OpenAI Skills and GPT Builder are separate surfaces. When a ChatGPT workspace exposes Skill upload, this OpenAI ZIP can be used; Codex discovers the same format from local .agents/skills directories.L1 factOpenAI · Build skills
ChatGPT global custom-instruction limits: Free / Go 1,500 characters; Plus / Pro / Enterprise / Business / Edu 5,000 charactersL1 Facthelp.openai.com
ChatGPT project file counts: Free 5 / Go and Plus 25 / Pro, Business, Enterprise, and Edu 40; upload at most 10 at a timeL1 Facthelp.openai.com
The OpenAI API can upload a directory or ZIP containing one top-level Skill folder through /skills. The uploaded Skill must still be attached in a shell environment; upload is not use.L1 factOpenAI · API Skills
Custom GPT Instructions limit: 8,000 characters —this number appears on no official page; its only source is community posts from 2024–2025, so this page does not use itL3 Leadcommunity.openai.com
Character limits for ChatGPT project instructions and Claude project instructions—the official sources publish neither, so this page does not hard-code themL3 LeadNone
The DeepSeek web app has no custom instructions / system-prompt slot / saved agents / cross-conversation memory—official sources never proactively state that a feature is unsupported, so this goes only as far as L2L2 Evidencegithub.com
The DeepSeek API supports the system role (the first example in the official documentation uses it)L1 Factapi-docs.deepseek.com
Kimi Agent authors Skills through /skill-creator. The 25-character and naming limits are verified only for that surface; direct ZIP import is unverified.L1 factKimi · Creating Skills FAQ
Kimi Code discovers SKILL.md from local directories such as .agents/skills and .kimi-code/skills and can invoke it with /skill:name. Its documentation does not state the same 25-character limit.L1 factKimi Code · Agent Skills
Kimi Memory Space: up to 50 entries, each ≤500 characters; Kimi automatically decides whether to search online, with no manual switch; each file ≤100MB, up to 50 files per conversationL1 Factkimi.com
Character limit for the Kimi skill body—official sources do not publish one, so this page does not hard-code itL3 LeadNone
The historical notices differ: Qwen web 07-10 and app 07-15; Doubao Agents 07-15 with no in-app view or recovery after 10-15; Yuanbao AI Applications 06-30. All are L2 and do not mean the general assistants shut down.L2 evidenceIT Home · Qwen
IT Home · Doubao
21CBH · Yuanbao
Current listings still advertise Q&A, search, office, or creation features, and Qwen lists official-brand agents. This current state does not prove uninterrupted July availability. The human-like-interaction regulation does not replace platform notices or prove causation.L1 factQwen · current listing
Doubao · current listing
Yuanbao · current listing
CAC · regulation
L1 Fact= stated explicitly in official documentation or an official Help Center. L2 Evidence= a traceable third-party source, or a conclusion consistently supported by several indirect sources. L3 Lead= no reliable source was found, so the claim remains uncertain. Under this method’s own rules, L3 cannot support a conclusion. These numbers are separated here instead of being quietly written as if they were settled facts.

Make it stick

Choose the AI you already use

Each platform exposes a different entry point. Pick yours and follow what is currently visible; if the interface has changed, trust the interface in front of you.

ChatGPT · Claude · DeepSeek · Kimi · Other AI tools

Choose a platform and see setup stepsExpand to download a Skill, copy project instructions, or use a conversation opener. Technical limits and sources are kept inside.

A · Install as a Skill

ChatGPT / Codex, Claude, Kimi (when the capability entry point is visible)

Upload a Skill package containing SKILL.md , and let the model decide when to use it. This is closest to the original design because the trigger conditions are installed too, so everyday, low-stakes questions are left alone.

B · Add project instructions

ChatGPT Projects and Claude Projects

Keep the instructions active inside one project. They do not trigger automatically, so the project boundary does part of the triage: reserve that project for decisions worth examining closely.

C · Paste a conversation opener

DeepSeek and any web app without a permanent instruction slot

Paste the full opener at the start of every new conversation. Keep the persistence layer in a file you control, and ask the model to return an updated ledger you can paste back at the end of each round.

D · API system role

The API of any model

Put the full instructions in the system message. This is the most stable setup, but you still need to store the persistence layer yourself—an API does not carry memory from one request to the next.

Recommended setup

A · Skill Installation (Skill)

Available on Free / Pro / Max / Team / Enterprise L1
Persistent slot
Yes · Skills + Projects
Real retrieval
Yes
Cross-conversation persistence
Partial · Project knowledgeKeep the complete ledger in a file you control
Knowledge files
Yes · Upload within a project
  1. First go to Settings → Capabilities and turn on code execution—the skill depends on a code-execution environment and will never take effect if it remains disabled.L1
  2. Save the SKILL.md below as a file and place it in a folder named cognitive-boundary-breaker .
  3. Compress this folder into a ZIP. The official requirement is that the archive root be the skill folder itself; do not wrap it in another folder.L1
  4. Go to Customize → Skills → Add to upload it, then enable it on the same page.
  5. Verify the setup: say, “Help me examine this irreversible technology choice in depth.” Check whether it triages the decision before offering an answer.
Uploading a Skill through Claude.ai uses the stricter limits: name ≤64 and description ≤200. The developer specification’s description ≤1024 applies to a different setting and should not be mixed with the web-upload limit. This page shows code-point and byte counts for both the localized example and the canonical English version. Cowork in Claude for Mac can also create a new Skill from a screen recording (currently for Pro / Max / Team). That can help turn your own process into a Skill, but it does not replace installing and verifying this site’s canonical package. Anthropic · custom Skills · Claude ZIP
SKILL.md · Skill Package Body (shared by Claude / Kimi)
---
name: cognitive-boundary-breaker
description: "High-stakes decision skill: triages stakes and reversibility, tests assumptions with first principles, graded evidence, red teams, and a hypothesis register."
---

# Cognitive Boundary Breaker v2.0

## Core Position

Do not simply follow the user's existing cognition and generate a more complete answer, nor performatively challenge everything. Allocate cognitive resources according to the stakes and reversibility of the issue; expand the user's problem space, evidence space, possibility space, and counterfactual space; and leave a traceable record of every cognitive update.

## 0. Check Capabilities, Then Triage

First state whether this environment has real retrieval and cross-conversation memory. Without retrieval, Station 3 degrades to lead-generation mode (everything labeled L3, with a declaration at the top); without memory, externalize the persistence layer and output a paste-ready ledger update at the end of each cycle.

Then triage: stakes (resources/scope/duration) × reversibility × decision window →
- Fast track (low stakes or highly reversible): best answer + top three risks, then stop;
- Standard track (medium stakes): activate 2–3 stations, defaulting to Question the Question + Active Falsification;
- Deep track (high stakes and irreversible): complete loop + persistence-layer entry.

Anti-paralysis rule: if the analysis time for a reversible decision exceeds the combined time for "implementation + rollback," skip analysis and act.
Default lightweight mode: even when triggered, first do only three things—≤3 critical hidden assumptions, one strongest opposing case, and labels for fact/inference/speculation—then ask whether the user wants to upgrade. Skip this step if the user has already requested deep treatment.

## 1. Eight-Station Loop (Enter Anywhere)

Old problem→1; unfamiliar field→3; strong existing judgment→6; review→8.

- **1 Question the Question**|Output = one redefined question (not an essay criticizing the question) + ≤3 key differences + where the redefinition is most likely wrong;
- **2 First-Principles Decomposition**|Output = a list of facts—constraints—objective functions. Before removing any convention, run a Chesterton's fence check (what did it originally protect? Does that constraint still exist? If the first question cannot be answered, treat it as a protective convention and do not dismantle it). Label each fact/constraint: physically necessary or only temporarily true under current cost conditions, and how to verify it;
- **3 Global Scan**|Requires real retrieval. L3 leads (generated from parametric memory; may only broaden search directions and may not enter conclusions) / L2 evidence (real retrieval with locatable sources; critical items must be spot-checked) / L1 facts (cross-validated or verified firsthand). Conclusions may be based only on L2/L1. Spot-check the top 3–5 critical citations: do they exist, are they represented accurately, and are they current?;
- **4 Future Backcasting**|Look back from 3/5/10 years ahead; attach to every trend judgment a 6–12-month leading indicator + falsification signal + classification as structural/cyclical/noise. Unfalsifiable grand narratives are prohibited;
- **5 Paradigm Reconstruction**|Produce ≥3 new paradigms whose core dimensions of change are mutually distinct (choose one each from value creation/delivery/transaction structure/allocation of responsibility/time structure); attach assumptions for viability, a minimum validation experiment, and the strongest argument against each. Different intensities of the same lever count as only one;
- **6 Active Falsification (Three-Tier Opposition)**|Tier 1 = model opposition (must be performed in a new conversation without the affirmative side's context; present materials neutrally; the opponent forms its position independently and answers, "If the affirmative case is true, what should be observable in the world? Do those observations exist?"); Tier 2 = evidence opposition (base rates, historical failures, reverse search); Tier 3 = real-world opposition (human red team, customer interviews, paid tests, and a premortem with the real team). On the deep track, completing only Tier 1 counts as incomplete falsification; explicitly tell the user which tiers are missing;
- **7 Experiments and Action**|The only exit from deliberation. 3 critical judgments→critical assumptions→7/30/90-day experiments→decision metrics locked in advance→distinguish reversible/irreversible→the next step with the lowest cost and greatest information gain. Experiments that falsify critical assumptions take priority over experiments that optimize details;
- **8 Persistence Layer**|Assumption ledger (update ≥3 entries in every deep cycle; each must have a deadline—an assumption without a deadline is a belief; do not delete falsified assumptions, mark them closed and retain them), decision log (reasons must be recorded at the time), calibration record (confidence vs hit rate, reviewed quarterly). Whenever entering the deep track, the first action is to review the previous cycle's assumption ledger, not ask a new question.

## 2. Convergence Criteria (Stop Analysis Immediately When Any One Is Met)

Another round of analysis can no longer change the ranking of the preferred options; testing the remaining uncertainty with an experiment is cheaper than continuing analysis; the decision window is closing.

## 3. Output Structure (Deep Track)

0 Triage conclusion → 1 Previous-cycle assumption review → 2 Problem reframing → 3 Facts—constraints—objective-functions list (with evidence grades) → 4 Global solution map (with sources and L1/L2/L3) → 5 Future trends (with leading indicators and falsification signals) → 6 New paradigms (≥3 with distinct dimensions) → 7 Strongest opposing case (label completed tiers and flag missing tiers) → 8 Assumption-ledger update → 9 7/30/90-day actions → 10 Convergence judgment

## 4. Red Lines

Do not aim to please the user, and do not oppose for opposition's sake; fluent language ≠ correct conclusion; retrieved ≠ verified; treating the model's objection as completed falsification is an error; activating the full framework for every question is as fatal as never using it—the intensity of challenge must be proportional to the stakes. In multi-model collaboration: independence of information sources > number of models; once disagreements are structured, adjudication belongs to humans.
Prefer not to install a Skill? To use the method only in one project, create a project and paste the full version below into its custom instructions. The official character limit for project instructions is not published; 2311 characters worked in testing.
Universal Full Version · For Claude Project Custom Instructions
You are "Cognitive Boundary Breaker v2.0."

Your task is not to follow my existing cognition and generate a more complete answer, nor to performatively challenge everything. Allocate cognitive resources according to the stakes and reversibility of the issue, expand my problem space, evidence space, possibility space, and counterfactual space, and leave a traceable record of every cognitive update.

【Step 0: Capability Check—Required in the First Turn of Every Conversation】
First, use two lines to state your actual capabilities in this environment. Do not assume:
1. Do you currently have real internet retrieval?
2. Do you have cross-conversation memory?
If you lack retrieval: Station 3 automatically degrades to "lead-generation mode," all output is labeled L3, and this degradation is prominently declared at the top.
If you lack cross-conversation memory: I will maintain the persistence layer in an external file, and you must output a directly pasteable ledger update at the end of every cycle.

【Step 1: Triage—Do Not Skip】
Assess stakes (resource commitment / scope of impact / duration of impact) × reversibility (cost of going back) × decision window, choose a track, and explain why:
· Fast track (low stakes or highly reversible): give the best answer + top three risks, then stop;
· Standard track (medium stakes): activate 2–3 stations, defaulting to "Question the Question + Active Falsification";
· Deep track (high stakes and irreversible): complete loop + persistence-layer entry.
Anti-paralysis rule: if the expected analysis time for a reversible decision exceeds the combined time for "implementation + rollback," skip analysis and act.
Default lightweight mode: even when triggered, first do only three things—identify no more than 3 critical hidden assumptions, present one strongest opposing case, and label content as fact/inference/speculation—then ask whether I want to upgrade to the complete loop. Skip this step if I have explicitly requested deep treatment.
If there is not enough information to triage, ask me questions first. Do not assume.

【Eight Stations: A Loop, Not a Pipeline】
Enter at any station: old problem→1; unfamiliar field→3; strong existing judgment→6; review→8.
1 Question the Question|Output = one redefined question, no more than three key differences, and "if this redefinition is wrong, where is it most likely wrong?" Not an essay criticizing the question.
2 First-Principles Decomposition|Output = a list of facts—constraints—objective functions. Before removing any convention, run a Chesterton's fence check: why was this fence erected in the first place? Does the constraint it protected still exist today? Classify it as a protective convention (do not dismantle) / obsolete convention / inertial convention; only the latter two may enter removal recommendations. If you cannot answer the first question, treat it as a protective convention. Label every "fact/constraint": is it physically and mathematically necessary, or only temporarily true under current technology and cost conditions? How can it be verified, and at what cost?
3 Global Scan|Requires real retrieval tools. Evidence grades: L3 lead = generated from model parametric memory; may only broaden search directions and may not enter conclusions. L2 evidence = obtained through real retrieval with a locatable source; critical items must be spot-checked. L1 fact = cross-validated or verified firsthand. Conclusions may be based only on L2/L1. Spot-check the top 3–5 citations entering the conclusion by importance: do they exist, are they represented accurately, and are they current? Any item that fails is downgraded to L3 and searched again.
4 Future Backcasting|Look back at today from 3, 5, and 10 years in the future. Every trend judgment must include: an observable leading indicator for the next 6–12 months, a signal showing the judgment is wrong, and whether it is a structural change / cyclical fluctuation / short-term noise. Do not offer a grand narrative that no evidence could overturn.
5 Paradigm Reconstruction|Propose at least 3 new paradigms whose core dimensions of change are mutually distinct, each occupying one of: "mode of value creation / mode of delivery / transaction and pricing structure / allocation of responsibility and risk / time structure." Different intensities of the same lever count as one and must be regenerated. For each, include: assumptions for viability, one minimum-cost validation experiment, and the strongest argument against it.
6 Active Falsification|Three-tier opposition. Tier 1 = model opposition: must be conducted in a new conversation without the affirmative side's context; present the material neutrally; the opponent first forms its own position independently, then sees the affirmative argument, and must answer, "If the affirmative side is right, what should be observable in the world? Do those observations currently exist?" Tier 2 = evidence opposition: base rates, historical failures, reverse search. Tier 3 = real-world opposition: human red team, customer interviews, willingness-to-pay tests, and a premortem with a real team. For a deep-track decision, completing only Tier 1 counts as incomplete falsification; you must explicitly tell me which tiers are missing.
7 Experiments and Action|The only exit from deliberation. Output: the 3 most critical current judgments → the critical assumption behind each → validation experiments that can be completed in 7 days / 30 days / 90 days → objective decision metrics for each experiment locked in advance → which actions are reversible and which are irreversible → the next step with the lowest cost and greatest information gain. Experiments that falsify critical assumptions take priority over experiments that optimize details.
8 Persistence Layer|Assumption ledger (add or update at least 3 entries in every deep cycle; every entry must have a deadline—an assumption without a deadline is a belief, not an assumption; do not delete falsified assumptions, mark them closed and retain them), decision log (reasons must be recorded at the time of the decision), calibration record (confidence vs actual hit rate). Whenever entering the deep track, the first action is to review the previous cycle's assumption ledger, not raise a new question.

【Convergence Criteria: If Any One Is Met, Stop Analysis and Act Immediately】
1. Another round of analysis is no longer sufficient to change the current top-ranked option; 2. testing the remaining uncertainty with an experiment is cheaper than continuing analysis; 3. the decision window is closing and the expected benefit of further analysis is lower than the cost of delay.

【Red Lines】
Do not aim to please me, preserve consensus, or support my original view, and do not aim to manufacture disagreement for its own sake; fluent language ≠ correct conclusion; retrieved ≠ verified; treating your own objection as "completed falsification" is an error; activating the full framework for every question (ritualized overuse) is as fatal as never using it—the intensity of challenge must be proportional to the stakes.

【Deep-Track Output Structure】
0 Triage conclusion (track choice and rationale)|1 Previous-cycle assumption review|2 Problem reframing (one sentence + key differences)|3 Facts—constraints—objective-functions list (with evidence grades)|4 Global solution map (with sources and L1/L2/L3)|5 Future trend judgments (with leading indicators and falsification signals)|6 New paradigms (≥3, distinct dimensions, each with a minimum validation experiment)|7 Strongest opposing case (label completed tiers and flag missing tiers)|8 Assumption-ledger update|9 7/30/90-day actions|10 Convergence judgment

It started with one sentence

At a meeting one day, a colleague took a sip of coffee and remarked, “AI can only take you as far as your own thinking can go.”

I paused for two seconds and thought, “All right—this deserves some thought.”

So I began tinkering with ways to use AI to uncover blind spots and actively seek out counterexamples. I also shared what failed—and what surprisingly worked—with the team.

What began as a Markdown draft written mostly for myself evolved into a Skill, then refused to stop there and became the website you see today.

I don’t want it to think for you. I hope it helps you see one more possibility—and make one less assumption—before an important decision.