Why Agent Feedback Stops Working and a Loss-Based Alternative: The Deep Twin Design
Who this is forBuilders of AI agents and no-code automation who keep patching skill files with feedback and want to understand why quality stops improving.
If you have built an AI agent with a skill file, you have probably watched its quality rise after each round of feedback and then fall again when you add a new test case. The usual fix is to add more instructions. This article reconstructs a 31-minute Korean video lecture on the “Deep Twin” design, which argues that the problem lies in the form the feedback takes. You will see the reported test sequence, why splitting one agent into several can compound errors, the loss-based feedback idea borrowed from deep learning, and a list of what the presenter has not yet revealed. The design is a hypothesis, not a verified result, and this article keeps that distinction throughout.
Loss-based feedback loop diagram

Core data
Video information
| Item | Value |
|---|---|
| Title | Almost working… then tearing it all down and starting over again | Presentation on the inherent limits of current AX approaches and a solution design |
| Channel | Yang’s Vibe Coding University (a Korean YouTube channel) |
| Upload | September 8, 2026, 31:10 runtime, 78,049 views (as of September 19, 2026) |
| Format | Whiteboard lecture and Part 1 of the Deep Twin series. The description box links to a Fast Campus course (a South Korean online education provider), an AX consultant training program, and 1:1 coaching, so the video mixes in promotion for paid education. AX refers to AI transformation, a common term in Korean industry |
| Captions | Korean auto-generated captions (generated by YouTube). No manual captions |
1. How agent refinement gets stuck (0:00–10:09)
- When you ask Codex or Claude to build an agent, the “100 out of 100” result is a single skill (2:46). A skill is a Markdown file with frontmatter (a name and when to use it) and a work set that runs from inputs to process to outputs. Instructions such as “do not do this” and “focus on this” pile up on top of it (3:10).
- The example is a proposal-writing agent (3:40). It takes four inputs: customer information, our service package, a proposal template, and your own comments, and it produces a proposal.
- The first test (Q1) reaches about 50% (5:10). The output differs from “the quality or approach I would have produced myself.” Feedback modifies the skill file and adds instructions, and the score climbs to 75% and then 90% (5:50).
- The next week, Q2 is run and yields 60% (6:20). Feedback brings it back to 90%, but once the instructions overload, rerunning Q1 drops to 80% (6:50). The presenter describes this as “poke one thing and another stops working.” The agent eventually freezes into something unusable (6:50).
- Rebuilding from scratch gives 80% from the start, which feels like growth. But a person with strong metacognition notices that the result is “not an agent, just more prompt added each time” (7:20).
- So the single agent is split into a multi-agent system (7:43). Research, drafting, review, and a pass/fail branch each become an agent, followed by graph engineering, loop engineering, and a control UI (8:30).
- Multi-agent systems hit the same problem (9:02). Even if each individual agent is raised to 95%, the final output is the product of those values. The presenter puts it this way: multiply 0.999 by 0.999 indefinitely and you get 0 (9:20). Concepts also multiply, until the person building the system no longer knows what the prompts say (9:50).
2. The core problem: explicit and tacit knowledge (10:09–15:06)
- The analogy begins at an exhibition (10:30). A girlfriend says “I like that painting” and gives a reason. If you write that reason down explicitly and apply it to other paintings, can you predict what she will like? The presenter says absolutely not (11:45).
- Explicit knowledge is judgment criteria that can be put into words, written down, and fully applied. Tacit knowledge sits beneath that and cannot be expressed in words (12:00).
- Back to proposals: the four inputs can be defined, but the process from inputs to output involves tens of thousands of judgments based on experience, such as research and review (12:40). Writing “this is how I judge” over and over has the same problem as “that’s why I like it” in art appreciation (13:20).
- The presenter argues that second brains and LLM wikis are also explicit-knowledge forms and cannot escape the same inherent limit (13:50). This is the presenter’s opinion.
- The conclusion (14:30): diligently giving feedback to keep an agent in line amounts to updating explicit knowledge with more explicit knowledge. That loop is inherently limited when it comes to drawing out your own tacit knowledge.
3. Clues from deep learning (15:06–21:10)
- The presenter originally worked in deep learning for image vision (15:30). Machine learning repeatedly trains on a dataset with labeled answers to build a model that predicts answers for new, similar data (16:20).
- The training process (17:00) works like this: hide the answer and have the model predict it. The degree of error is the loss. Then iterate, searching for weights that reduce the loss. Each iteration is an epoch. Plot epochs on the horizontal axis and accuracy on the vertical axis, and the curve starts near 0, rises, flattens, and eventually stops moving. Training ends at that plateau, and the saved weights are the model (18:00–19:20).
- The clue (19:51): the algorithm lowers the loss, but people cannot tell how it adjusted the weights. Nobody gave explicit feedback such as “bend it 270 degrees here.” When you open the model file, it contains no rules, only numbers whose meaning you cannot read, yet it gets the answers right (20:30). The presenter asks whether it is fair to imagine that this model file holds deep tacit knowledge (21:00).
4. Flipping the feedback method (21:10–22:46)
- Until now, you looked at a result and gave explicit feedback about how it should have been done (21:30).
- Flipped, the approach works like this: when a result comes out, you give no “change it this way” statement. You provide only your own version of the output you would have produced (21:50).
- The gap between the agent’s output and yours can then be treated as a loss (22:00). The presenter offers two hypotheses (22:20–22:40). First, reducing this loss is not explicit knowledge, so something must sit in the place of the algorithm in deep learning. Second, as long as it is not explicit knowledge, whatever fills that place should work.
5. How the Deep Twin tool works (22:46–28:16, presented as an imagined tool)
- The user attaches files to a prompt, turns on the microphone, and talks for about 30 minutes about their work process (23:10).
- The tool listens and draws the work process as a graph, then asks, “Is it like this?” The user does not need to know that it is a graph (23:50). After feedback, graph 0 is finalized (24:20).
- You feed in a first task, Q. Result R1 is just as bad as what you would get without the tool (24:40).
- The feedback is a comparison: “how I would have made R1” (25:00).
- The tool calculates the loss, the difference between R1 and your feedback, and changes the graph to reduce it. Inside the graph are memory in each node (agent), propagation across edges, and a dense design of judgment grounds. People cannot read the interior the way they would read a model file (25:30–26:00). The modified version is graph 1-1.
- Feeding the same Q into graph 1-1 produces R1-1. That output is compared against the original feedback to get loss 1-1 (26:20).
- The process repeats until the curve flattens and there is no sign of further improvement (26:50). The final graph is stored in the tool. The process produces what corresponds to the model in deep learning, which the presenter calls Deep Twin’s tacit knowledge (27:10).
- The presenter did not reveal what that tacit knowledge looks like. Is it Markdown, which would make it explicit again, embedding numbers, or an attention concept? (27:30–28:10). The presenter judged that it should be revealed only after the tool is developed and validated.
6. Publication plan (28:16–31:10)
- The imagined tool is being developed as a real open-source framework. It is a starter kit that lets users extend it freely, as long as the process flow is preserved (28:40).
- Planned releases include one foundational paper, a tutorial, documentation, and third-party validation results (29:00–29:40). The actual shape of the tacit knowledge will be revealed with the tutorial.
- The presenter’s confession (30:20): it has been hard to explain the Deep Twin concept so far because the substance of the core design was lacking even for the presenter. The earlier videos will not be deleted.
Insights for BuildnWrite
We selected the parts of this video that are useful for BuildnWrite, which covers content, courses, and automation.
- Our harness shows the same symptoms. Rules pile up in
harness/rules-common/and the project CLAUDE.md, and we are considering a rule diet, which we put on hold on September 3, 2026. That is the “instruction overload, then poke one and another breaks” pattern from the video. The presenter’s diagnosis, updating explicit knowledge with explicit knowledge in a loop that never ends, points the same direction as the line inbehavior-interactive.mdthat asks whether removing rules would be better before adding one. When we teach why adding rules can lower quality, the 50→90→60 storyline from this video makes a good opener. - Writing feedback as “here is what I would have produced” can be done today, without the tool. Our reviewer agents (deck-reviewer, workbook-reviewer, and others) return verdicts, and we turn those verdicts into rules. Keeping the user’s revised final version, which serves as the correct answer, next to the draft to show the difference is already partly done in our SNS writing-style feedback (similar to
feedback_linkedin_writing_style.md). Systematizing this as “keeping answer pairs” becomes example-based few-shot prompting, even without Deep Twin. The video goes further than using examples instead of rules, though. It claims the graph itself gets updated. What we can do is the part before that. - The multiplication argument can be cited as a counterpoint to multi-agent design. The arithmetic of 0.95 raised to the power n heading toward 0 connects to why our research skill gates the deep-research multi-agent setup to high-risk investigations only. However, this logic assumes that the steps are independent and that errors accumulate. If a review stage corrects earlier errors, the result is not multiplication. Attach that caveat when citing it.
- Value and limits as content. With 78,049 views, the topic has traction in the Korean vibe-coding community. The phrase “turning tacit knowledge into assets” overlaps with the readers of our book, a Korean-language title roughly translated as “The Art of Making Claude Work Hard.” However, this video is a design hypothesis, not verified results, and its core, the storage format of tacit knowledge, is undisclosed. Our courses and writing should not present Deep Twin as “this is how it works.” The most we can say is that someone has proposed this approach. When the open-source project or paper is released, we will update this with follow-up research.
- Points to critique. (a) The failure of explicit feedback is explained only as skill-file bloat. It can also be explained as overfitting from tuning against a single test case. In that case the fix is to expand the test set, not to adopt Deep Twin. (b) The deep-learning analogy assumes many correct-answer pairs and a quantifiable loss. The video does not explain how to quantify the difference between two documents in outputs such as proposals. (c) Calling second brains and LLM wikis limited because they are explicit knowledge is the presenter’s opinion, and it conflicts with our wiki-ingest feedback loop. Adopting that view would require stronger evidence.
Quotable lines (short quotes are translated from the source; everything else is summarized):
- “Multiply 0.999 by 0.999 indefinitely and you get 0” (9:20)
- The presenter’s remark that the agent becomes “just more prompt added each time” (7:30)
- Summary: give feedback as “what I would have produced” instead of “change it this way,” and treat the gap as loss (21:50–22:10)
- Summary: a model file with no readable rules, only numbers whose meaning is unknown, gets answers right, which suggests it may hold tacit knowledge (20:30–21:00)
Source reliability
All facts in this article were checked on September 19, 2026, and every factual claim comes from the captions below.
| Source | Type | How verified | Reliability | Where used |
|---|---|---|---|---|
| Yang’s Vibe Coding University, “Almost working… then tearing it all down and starting over again” (September 8, 2026, 31:10) | YouTube lecture. The presenter runs a paid education course and previews their own framework, so the view favors their own method | Metadata, chapters, description, and full Korean auto-caption transcript retrieved with yt-dlp. Originals: transcript-full.txt, transcript-full.ko.vtt |
Medium. The statements are confirmed, but the auto-captions contain errors. Several words were corrected by hand, including the garbled terms for tacit knowledge, Deep Twin, proposal, and result. The percentages (such as 50% and 90%) are qualitative examples, and the presenter said there are no quantitative figures (6:00) | All sections |
| Video description: timeline and “what you will learn in this video” | Summary written by the channel | yt-dlp description |
High, since the channel wrote it. Checked against the captions and matched | Video information, section breaks |
Unverified: As of September 19, 2026, the presenter had only said that the Deep Twin open-source project and paper would be released later. We did not find either in a separate search.
Bottom line
The evidence supports a narrower conclusion than the video does. Accumulating rules can make an agent worse, and splitting work across agents can compound errors when no stage corrects earlier ones. Feedback that shows the output a person would have produced is a useful way to build answer pairs, and that is worth doing now. The claim that a loss-driven graph update produces reusable tacit knowledge remains a hypothesis. The storage format is undisclosed, no third-party results exist yet, and the deep-learning analogy assumes a quantifiable loss that the video does not establish. Treat Deep Twin as a proposal to watch, not a method to teach.
Sources
- https://www.youtube.com/watch?v=wRgJPO_qIT4
- ./assets/2026-09-19-deep-twin-tacit-knowledge-agent/transcript-full.txt
- ./assets/2026-09-19-deep-twin-tacit-knowledge-agent/transcript-full.ko.vtt
- ./assets/2026-09-19-deep-twin-tacit-knowledge-agent/diagram-loss-feedback-loop.html
- ./assets/2026-09-19-deep-twin-tacit-knowledge-agent/diagram-loss-feedback-loop.png
Frequently asked questions
- Why does adding more feedback rules to an agent eventually hurt its results?
- The presenter describes a test where rule additions raised a first task to 90%, but a second task dropped to 60%. Rerunning the first task then fell to 80% because accumulated instructions overload the agent.
- What does the Deep Twin design change about feedback?
- Instead of saying how to change an output, the reviewer supplies their own version of it. The gap becomes a loss that the tool reduces by updating the agent graph. The design is an unverified hypothesis.
BuildnWrite helps teams build AI agents that keep running. About BuildnWrite ›