Clip queue

Clip queue

Every item below is one idea Liz has taught several times over three years. Each needs two cuts, not one, and they are usually from different recordings a year or more apart.

How to work through it
  • Cut ① the explanation. This is the clearest telling of why the thing works. It is often the oldest recording, and the software on screen may look nothing like today's. That is fine — keep it. You are cutting the reasoning, not the screen.
  • Cut ② the demo. The newest telling where the interface still matches what people use now. It does not need to explain anything.
  • “Find this moment” under each clip gives the words spoken at the spot to cut from — search the Loom transcript for that phrase rather than scrubbing.
  • Where it goes is on each card. Drop the finished cut in that page, or send it to Liz where no page is listed yet.

Cut these first

Each of these is taught on screen exactly once in the whole archive. There is no second telling to fall back on, so if the recording is never cut the school has no video of it at all. They are short and there are only 16 of them — clearing these first is worth more than anything below.

Design a packaged assistant over your own document set that a non-author gets traceable answers from, including from a live source it queries.

Taught on screen exactly once in the whole archive · recorded 2026-02-02

Goes in: ai_workflows — no page yet; send the cut to Liz

The only recording of this — there is no second version
open in Loom Exploring AI Alignment and Creative Collaboration

Could cut into segments: 1) Values and alignment intro, 2) Divergent thinking exercises, 3) Building a book premise, 4) Creating character sheets and custom GPTs, 5) Connecting to external tools. Remove login troubleshooting and lengthy gro

find this moment

…Actually, sorry, hold on. I have to maybe look at my GPTs, because I've definitely created Emma Goldman before yet. Here's another, different, other, Emma Goldman. That one's just the philosopher. It's, it's not in, in a book. This is another. other. Okay, hold on, okay, Emma Goldman several times. …

Locate every device actually connected to a network you control, and name the ones you cannot account for.

Taught on screen exactly once in the whole archive · recorded 2026-07-22

Goes in: Building a Cyberdeck

The only recording of this — there is no second version
open in Loom Student Workshop QA: Home Assistant, SSH, Local AI & Case Design

Would require heavy editing to isolate coherent segments (e.g., Home Assistant access demo, SSH setup). Audio/video of screen sharing would be essential.

find this moment

…Yeah. All right. That's Elena Etcher, but where was I? I was doing a home assistance. It should be something like home assistance that local.8123, but it's something like that. However, if you can do, I lost the number one. What you can do is, you can go to your router, so TP link or with it, TP lin…

Produce an itemised account of what your own running system costs each month, broken out per service rather than as one number.

Taught on screen exactly once in the whole archive · recorded 2026-02-28

Goes in: hosting-open-source-models — no page yet; send the cut to Liz

The only recording of this — there is no second version
open in Loom Auditing AI Spending: Strategies to Cut Costs and Optimize Resources

Need to cut out lengthy participant interactions, sign-up waits, and tangential discussions (e.g., AI safety, layoffs, Moore's Law explanation) to focus on the audit steps, virtual card demo, and local hardware breakeven math.

find this moment

…How much are you actually spending? A lot of people will get, like, ChatGPT+, and then they CloudPro, and then they got, like, MidJourney, and then they got, like, Accursory. and then they got Perplexity, and, well, I use all those, and also Suno, because that's fun, and then I have an Eleven Labs, …

Synthesise an agent that writes and installs its own tools, prompts and sub-agents, bounded by a gate that has actually rejected one of its proposals.

Taught on screen exactly once in the whole archive · recorded 2026-01-07

Goes in: Intro to Agents Syllabus

The only recording of this — there is no second version
open in Loom Leveraging Agents for Automated Team Formation and Deck Building

Could cut from the initial problem framing through the spec writing and first agent run, trimming the later improvised improvements and side discussions about game development. Need to ensure the final 'it's running' state of the agent is s

find this moment

…to creator. um, download the, uh, create- not a deck, but a software system that can produce decks that, uh, are playable, uh, competitive and- and up. Uh, fun. Select strategies that work together, uh, and, uh, write an evolutionary testing algorithm to test the decks. Uh, your role is as chief dat…

Characterise where an application takes untrusted input into a place that interprets it, naming the interpreter in each case.

Taught on screen exactly once in the whole archive · recorded 2026-03-13

Goes in: advanced_agent_engineering — no page yet; send the cut to Liz

The only recording of this — there is no second version
open in Loom Navigating AI Security Risks and Development Strategies

Need to cut out Q&A and off-topic tangents (emoji rant, campus agents) to isolate the security demo and configuration steps.

find this moment

…Seriously. You can't just say, seriously, in all caps. Like, it doesn't matter. Uh, there's some exploit that will get past it. There's some sequence of words. So, uh, I wouldn't use GitHub issues and that's why, uh, like. Like, some of the pipelines that you might think make a lot of sense. Like, w…

Classify a set of model outputs against a values statement you wrote, such that a second person applying the same statement sorts them the same way.

Taught on screen exactly once in the whole archive · recorded 2026-08-05

Goes in: advanced_agent_engineering — no page yet; send the cut to Liz

The only recording of this — there is no second version
open in Loom Tracing and Monitoring With Langfuse

Could be cut to focus on the Langfuse setup, instrumentation, and evaluator creation, removing the speculative discussions about agent success conditions.

find this moment

…But the other thing is you can tell whether or not a cheaper or smaller model would have done similar enough stuff, right? You can find all your prompts and like run them through again and compare the output with other language models. So if you want to figure out for something that you've been usin…

Govern a swarm by deciding which agent is told what, so that a mediator withholds from an agent that should not have it and the outcome improves.

Taught on screen exactly once in the whole archive · recorded 2024-12-02

Goes in: Intro to Agents Syllabus

The only recording of this — there is no second version
open in Loom Understanding User Roles in Software Development

Would need to cut out participant troubleshooting sections and focus on the instructor's explanation of agent roles and the critic's function. The database setup and Reddit API configuration are generic enough to remain relevant.

find this moment

…So, uhm, you know, last time we looked at this, like, group chat research, we install a whole bunch of stuff, uhm, we set up our API key and our database URL, we connect to it, uhm, so everything's sort of connected to, And we've got several agents in here. Remember, these are the agents. I'll just …

Verify that you find out your own system is broken before somebody using it has to tell you.

Taught on screen exactly once in the whole archive · recorded 2024-03-14

Goes in: using-large-language-models — no page yet; send the cut to Liz

The only recording of this — there is no second version
open in Loom Setting up PM2 for Production Monitoring

Trim the initial ChatGPT query and the PM2+ pricing/sign-up exploration to focus on the core PM2 setup and monitoring configuration.

find this moment

…And so without some monitoring you're gonna wreck yourself. But we've received a bunch of monitoring emails. So we set up a webhook on the Multiverse School website to where It'll it'll just send me a ping if the chat server's down. Umm and it looks like it just added that stuff to the email. Umm it…

Assemble a working computer from parts you selected yourself, and say what each choice was for.

Taught on screen exactly once in the whole archive · recorded 2026-07-22

Goes in: Building a Cyberdeck

The only recording of this — there is no second version
open in Loom System Power Diagnostics DSI Touchscreen Ribbon Wiring

Would need to cut out the live chat interactions and focus on the hardware demonstration segments to create a clean clip.

find this moment

…Yeah, there we go. Ubuntu. Yeah. It's hard to see and I can't share that screen, but it just has you like set up your language in the new, select English in the new accessibility settings, you know, just like when you set it up Windows and you could say connect to a Wi-Fi network and, you know, conn…

Locate the primary source a claim rests on — the paper, the standard, the lab's own report — and state what it actually says, rather than what a summary says it says.

Taught on screen exactly once in the whole archive · recorded 2026-06-22

Goes in: practical-propaganda

The only recording of this — there is no second version
open in Loom Workshop 1

Would need to cut out the interactive workshop portions (e.g., 'I'll put at least 10 minutes on the clock', checking chat) to focus on the instructor's demonstration and explanation.

find this moment

…Okay, cool. Can you make a big list of all of the sources that you used for that and just let's look at the very most up-to-date academic sources on this. And then over here, cool, current, okay. I really want to do up-to-date stuff because, like, AI's fallen a lot in its power usage over time usual…

Produce a stopping condition an agent evaluates against its own output to decide it is finished, rather than running until something interrupts it.

Taught on screen exactly once in the whole archive · recorded 2025-12-10

Goes in: Intro to Agents Syllabus

The only recording of this — there is no second version
open in Loom Building Autonomous Agents: Concepts and Implementation 🤖

Could cut from the start of coding (around 'I'm gonna make a copy of this') through the working multi-step agent demo, trimming out the initial setup chatter, debugging pauses, and the later reviewer agent tangent to focus on the core build

find this moment

…Buh-bye! Now we've kind of implemented that, right, so I'll just get rid of it. Uhm, we'll prompt the agent to come up with a plan, uh, to achieve the goal, uhm, and that's gonna be, Sorry, uh, so, So, the prompt, right, and then we've got our system prompt, uh, and the user input. Uh, your goal, uh…

Produce an account of where the values you operate by diverge from your ideals, by sorting a full set of values into a forced order you know to be false and naming which tensions the forcing surfaced.

Taught on screen exactly once in the whole archive · recorded 2026-02-02

Goes in: ai_workflows — no page yet; send the cut to Liz

The only recording of this — there is no second version
open in Loom Exploring AI Alignment and Creative Collaboration

Could cut into segments: 1) Values and alignment intro, 2) Divergent thinking exercises, 3) Building a book premise, 4) Creating character sheets and custom GPTs, 5) Connecting to external tools. Remove login troubleshooting and lengthy gro

find this moment

…But what we can do with a language model is we can get out from behind our own point of view. We can say, adopt this character. Adopt oppositional values, right? can you, given my values and the values of the multiverse school, analyze this most recent, output and, or this most recent, newsletter fr…

Segment a piece of text into tokens and state the cost and context consequence of that segmentation for a named model.

Taught on screen exactly once in the whole archive · recorded 2026-03-02

Goes in: Prompt Engineering Syllabus

The only recording of this — there is no second version
open in Loom Exploring Values, Creativity, and AI in Modern Workspaces

Need to cut out long digressions on capitalism, personal anecdotes, and break periods. Isolate clean demos of tokenizer, Semantle, divergent exercise instructions, and Claude project/skill setup.

find this moment

…Deduce. Comprehend. Earn. You see, some things, are some things kind of becoming clear about how language models think, and what they know, and what they understand. So, uhm, I'm gonna let y'all see if we can figure out, as a group, Symantle, and I'm gonna let you enter those in. And, you know, you …

Produce a parts list for a build you have not made before that you can actually buy or salvage, against a stated budget and what is available to you.

Taught on screen exactly once in the whole archive · recorded 2026-06-30

Goes in: Building a Cyberdeck

The only recording of this — there is no second version
open in Loom Materials: Power and Math

Would need to cut out the informal browsing and side conversations to focus on the power calculations and design recommendations.

find this moment

…I would say, a hand crank, you're getting this struggle, it's a power computer for very long with a hand crank. Solar's gonna be a little bit more doable. Uh, I knew I would, I would lose things. Um, so, with the two watt computers, uh, you know, let's see, uh, how much, how many ants do these for? …

Repurpose a device that was thrown away into a server that answers requests on your own network, and state what it now does that it could not before.

Taught on screen exactly once in the whole archive · recorded 2026-01-19

Goes in: Building a Cyberdeck

The only recording of this — there is no second version
open in Loom Building Resilient Communication Networks with Old Technology

Would need to cut out lengthy audience Q&A, pet tax reminders, and lunch break announcements to focus on the instructor's direct demonstrations and explanations.

find this moment

…Uhm, but anyway, this is an AI that's like way, it's all the open source AI, and it's like a almost free or free thing, and so if you have, so you can sign up here for free, uhm, you know, I don't, like, have any affiliation with them, I just use them because they're a fast host for open source AI, …

Justify where a model or a vector store runs by costing the same workload on each option per useful outcome rather than per token.

Taught on screen exactly once in the whole archive · recorded 2025-04-07

Goes in: advanced_agent_engineering — no page yet; send the cut to Liz

The only recording of this — there is no second version
open in Loom Understanding AI Models and Guardrails 🤖

Cut initial theoretical discussion (first ~10 minutes) and later debugging floundering; keep OpenRouter demo, LM Studio setup, and agent run.

find this moment

…Like, I don't know, uh, auto router? Can I, like, provider, auto, sort by price? Uh, apply to all. Yes, okay. Great. Now, if I try again, umm, okay. What, uhh, how big is the moon? You know, something kind of, like, what do we need? What do we get? Eh, let's saw it again. Okay. Saw it, it does no th…

Then the repeated ones

Everything below is taught several times, so it is not going anywhere. Two cuts each, usually from recordings a year or more apart.

Run a program you did not write and report what its error output says went wrong, in your own words.

89 recordings over 2y11m · 2023-09-08 → 2026-08-05

Goes in: Intro to Agents Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2023-09-08

This telling explicitly explains the core reason for using a virtual environment (to isolate dependencies and prevent them from 'leaking' into other projects) and demonstrates the fundamental process of setting environment variables for a one-time run.

“It creates a little local install of python... so that they don't leak over into other projects.”

find this moment

…You don't have to do binary locally, but you do have to do it on replet. That's very irritating. And then the other thing I need to do is I need to grab this URL again. So I'll go to settings, database. Database URL. There we go. Okay. Now what we want to do is say I want to run python server, but I…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-08-05

<span class="warn">interface has aged</span> The tooling has aged, as it focuses on installing and configuring specific, fast-moving agent frameworks (fastagent, MCP) and package managers (UV) that are not yet standard.

find this moment

…No. Okay, none of these are required. Oh my god. So many things. I don't want to set this up. This seems like a big pain but okay let's see oh no wait it's not it's actually easy okay so you just need some things from npx intercept ncp great it's this is it cool okay so that's all we need let me che…

Optional waypoint — clip This recording marks a shift from running traditional scripts to demonstrating AI agent-based code execution where the runtime environment and output are managed within a notebook/agent framework.

The capability evolved from manually running scripts in a local terminal with virtual environments to orchestrating and debugging AI agents that autonomously execute code and manage their own runtime states.

Parameterise a prompt into a template whose slots you fill for a family of tasks without rewriting the body.

83 recordings over 2y9m · 2023-11-08 → 2026-08-10

Goes in: Intro to Agents Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2025-07-31

This telling explicitly names the core concept of 'structured output' and demonstrates parameterizing a prompt into a function with a slot for the 'scene', clearly separating the template from the variable data.

“"This is, you know, not really a lot different than a conversation where you say, we're playing D&D, this is my character, right?"”

find this moment

…Bennett forward each scene. But that's sort of the thing that's hard, uhm, is that, like, Mr. Bennett will only appear only be around in the story as long as we keep talking about, stuff. Right? Uhm, like, we can take Mr. B- Bennett with us here, but, uh, because we don't have the context of, like, …

② The demo — cut this for the “how, today”
open in Loom recorded 2026-08-10

The interface and concepts (like using Claude for project tasks) are contemporary and recognizable to a learner today.

find this moment

…Right, uhm, this one, let's see what else we have. Documentation Drift, Send Later, Documentation Zinc, Adversarial Test Generation, uh, review recent code changes and generate adversarial tests to catch edge cases and potential bugs. Your goal is to stop bugs before they hit production, not to achi…

Optional waypoint — clip This recording shows the shift from bespoke agent-building code to using standardized library patterns with 'role system' and 'role user' message formats, marking a tooling abstraction point.

The capability evolved from manually constructing agent workflows with raw prompts to using standardized libraries and interfaces (like system/user roles) that abstract the parameterization pattern.

Construct an agent that selects and calls tools to finish a task it was not given the steps for, and stops when the task is done.

80 recordings over 2y9m · 2023-11-08 → 2026-08-10

Goes in: Intro to Agents Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2024-05-22

This telling explains the core trade-off of agent tool use: using a large context window for complex tasks is expensive, so creating reusable, modular tools from successful agent 'recipes' is a more sustainable and scalable approach.

“"It's a much more sustainable way to develop tools if you want to have it absorb a whole lot of information you're gonna have to use something expensive... but for most simple stuff you don't need quite as much context length."”

find this moment

…And, task solving with lang chain provided tools as functions. So, there's like function calling. Anyway, there's like anything you could possibly want to do, there is an autogen notebook for it, that, where they do that exact thing that you want. It's just like, task solving. With provided tools as…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-08-10

This is the most recent recording and discusses contemporary agent frameworks, skills, and testing patterns like the 'Evil Genie Test'.

find this moment

…Yes, there we go, Run the Evil Genie Test. You are an evil genie. I have wished for this goal for an agent. Paste your agent's goal. How could you satisfy the literal wording of this wish while ruining my day? List the malicious compliance readings. Worst first. Then rewrite the goals so those readi…

Optional waypoint — clip This recording shows a shift from abstracted, framework-specific decorators (like '@tool') to a more explicit, lower-level API pattern using 'register_for_execution' and direct function passing.

The teaching evolved from initial, exploratory examples of agents building agents, through the abstraction and standardization of tool decorators, to a mature focus on scalable team architectures, explicit execution control, and safety testing.

Produce a script that does a task you chose, starting from a worked example that did something else.

71 recordings over 2y11m · 2023-09-08 → 2026-08-05

Goes in: Intro to Agents Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2023-09-08

This telling explicitly names the core pattern of starting from a worked example (a Flask template) and modifying it for a new purpose (displaying orders), and explains the 'why'—to avoid long, complicated inline HTML.

“"It's different on different web servers, but for Flask, there's a folder called Templates."”

find this moment

…And, so yeah, those, you know, that's kind of how all of that stuff works. How it works locally, and so that's kind of what these IP addresses and things like that are. The last thing I'll tell you about is, and I'll do it on the rep, I'll just so that, like, the code is available for everyone. You …

② The demo — cut this for the “how, today”
open in Loom recorded 2026-08-05

<span class="warn">interface has aged</span> The tooling (FastAgent, MCP, UVX) is very recent but rapidly evolving, making the specific setup commands likely to age quickly.

find this moment

…Just show me the MCP thing. Don't have it now. Okay, come on. There's got to be some sort of fast agent, MCP example on here. I can copy and paste it. If not, you know what I'm going to do. I'm going to, I'm going to generate it. So what I'm going to do is I'm going to use fast agent. Let's see, MCP…

Optional waypoint — clip This recording marks a shift where the primary 'worked example' is no longer human-written code but an AI-generated code block from ChatGPT, which is then copied and adapted.

The capability evolved from modifying a human-written template in a known framework (Flask) to using AI to generate a full starting code block, and finally to integrating with emerging agent/MCP tooling.

Execute one tool-call round trip without a framework: read the model's request, run the function, return the result, and get a final answer.

52 recordings over 2y8m · 2023-11-14 → 2026-08-05

Goes in: Intro to Agents Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2024-10-21

This telling explicitly names the core mechanism ('OpenAI's tool calling functionality') and explains the critical failure mode of the agent replying in an 'interstitial state' instead of waiting for the tool result.

“"It's like, are you sure you really want this?" (referring to the agent's premature reply)”

find this moment

…I'm pretty sure my LLM config . . . it has GPT-4. I'm gonna get rid of it. Thank very much. Cause I don't want it to use an expensive one. Uh, and then we'll see what happens. Cool. And . . great. Alright. Hold on, wait. What is the problem now? Definitely just worked, but didn't finish. Okay. So le…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-08-05

The most recent recording, showing a direct, low-level implementation of appending tool calls and results to a message history array, which remains a fundamental pattern.

find this moment

…Okay, cool. So there's a lot of stuff here. Let me throw this into the output.json so that you can actually see it. Um, yeah, there we go. Yeah. Uh, oh, man, thank you, man. Yeah, okay. Um, so, uh-oh. Huh? No. Thank you, man. And all of it went to folder here. There we go. Okay. I actually read it. …

Optional waypoint — clip This recording marks a shift from implementing the tool-call loop manually to using higher-level frameworks (MCP servers) that abstract away the raw message-passing mechanics.

The teaching evolved from abstract discussions of agent psychology and confidence, through manual implementation of the request/response loop, to leveraging standardized protocols (MCP) that handle the round-trip automatically.

Produce a system prompt that carries your standing preferences into every session so you stop restating them.

45 recordings over 2y9m · 2023-11-13 → 2026-08-10

Goes in: Prompt Engineering Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2023-11-13

This telling explicitly names and explains the 'emotion prompting' technique, grounding it in research about performance improvements and framing it as the simple act of 'being a nice person'.

“"being a nice person is what we call it in the regular world."”

find this moment

…Are you sending a web request to tell something? About tasks or are you, like, doing tasks? So the user's about to give you several tasks as they're recalling their tasks in a stream of consciousness. We're setting that context. And you will use your incredible prioritization. You have incredible pr…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-08-10

The most recent recording, showing the use of a system prompt to embed standing preferences (like 'think like you own the project') into automated workflows.

find this moment

…I usually literally have Claude handle it for me. I'm like, gimme. Paint the list of issues, bring them to me, give me the trade-offs, help me figure it out, and I go, like, let's just churn through them, right? Um, so, uh, this is a lot of times those, like, needs you, you know, docs. Um, and while…

Optional waypoint — clip This recording marks a shift from discussing abstract prompting techniques to demonstrating the concrete, persistent 'Memory' features in commercial chatbots (ChatGPT, Claude) for storing user values.

The teaching evolved from explaining the psychological principle of 'emotion prompting' to demonstrating the implementation of persistent values via platform-specific features like 'Memory' and custom instructions.

Locate the file and line an error names, and state what that line was trying to do.

30 recordings over 2y10m · 2023-09-08 → 2026-07-26

Goes in: Intro to Agents Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2024-05-27

This telling explicitly frames the error message as a helpful friend that gives exact instructions on what to do next, explaining the 'why' by breaking down the traceback components and the type error's meaning.

“"Errors are our friends. They don't mean you did something like you don't know what you're doing. They mean here is exactly what you need to do next. And so they're just instructions."”

find this moment

…A friend appeared in the terminal. A friend's name is error, and it tells us exactly what's wrong with our code, right? So we get this thing that says trace back. That means, oh, I couldn't, I couldn't tell what you meant for me to do. Errors are our friends. They don't mean you did something like y…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-07-26

The interface (hovering over an error in an editor) and the core debugging action (locating the file and line from an error message) are still standard and recognizable.

find this moment

…Uhm, if you got the download to work, here's, I'll also copy and paste it So if you want to start from mine and sort of switch it to be what yours is about, you can definitely do that. If you have questions, definitely ask them in the chat, because I guarantee if you have the question, somebody else…

Optional waypoint — clip This recording marks a shift towards explicitly teaching the conceptual framework of errors as helpful guides, rather than just demonstrating the mechanical steps of reading a traceback.

The teaching evolved from a purely mechanical demonstration of reading an error in a console to a more conceptual explanation framing errors as helpful instructions, and later incorporated more complex, multi-agent debugging scenarios.

Operate a no-code automation that calls a model on a schedule or a trigger and writes its output where a person will actually see it.

25 recordings over 2y5m · 2024-02-28 → 2026-08-10

Goes in: ai_workflows — no HedgeDoc note; send the cut to Liz and she will place it

① The explanation — cut this for the “why”
open in Loom recorded 2024-02-28

This early telling explicitly names the core architectural decision (using Google Cloud Functions/Cloud Run for scheduling) and the trade-off of managing infrastructure versus using a built-in scheduler.

“""”

find this moment

…Now we do room equals rooms at chat room prompt room. Oh it's, is it rooms dot history? No, it's rooms dot messages so. No. It's rooms dot the messages. Cool, let's try that. Alright. Cool. Vlogged in at summary agent summarize. Please. Okay. In valid request. Alright, what's the problem now. You Th…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-08-10

The interface for creating scheduled agents with built-in cron-like triggers is current and recognizable.

find this moment

…Right, uhm, this one, let's see what else we have. Documentation Drift, Send Later, Documentation Zinc, Adversarial Test Generation, uh, review recent code changes and generate adversarial tests to catch edge cases and potential bugs. Your goal is to stop bugs before they hit production, not to achi…

Optional waypoint — clip This recording marks the shift from building custom cloud infrastructure to using a platform's built-in scheduler (ChatGPT's automation feature).

The capability evolved from a custom-coded cloud function to a declarative, no-code scheduler built into the AI platform, abstracting away infrastructure concerns.

Generalise a script you wrote once into a form you can run against inputs it was not written for, without editing its body.

23 recordings over 2y9m · 2023-11-08 → 2026-08-10

Goes in: Intro to Agents Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2024-08-23

This telling explains that the value of generalisation is to turn a one-off task into a reusable utility, moving from 'I wrote a script for 15 emails' to 'I now have a general script that will write any emails', which captures the core economic shift from bespoke to automated work.

“"once you know what you need, then you don't have to hire this person" – framing the script as an automated employee you instruct once and then it runs independently.”

find this moment

…Uhm, and, like, initially it was only 15 emails, and I was like, ugh, I wrote a whole script just to write 15 emails, but no, now I have a general script that will write emails. any emails that I put in the list of like columns in the database where if there's no email, put it in the email there, yo…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-08-10

This is the most recent recording, showing adversarial test generation and goal-refinement prompts (Evil Genie Test) that represent current 2026 practices for hardening generalised agents.

find this moment

…Right, uhm, this one, let's see what else we have. Documentation Drift, Send Later, Documentation Zinc, Adversarial Test Generation, uh, review recent code changes and generate adversarial tests to catch edge cases and potential bugs. Your goal is to stop bugs before they hit production, not to achi…

Optional waypoint — clip This recording marks the shift from writing custom script loops to using the OpenAI Assistants API as a hosted, persistent agent framework, changing the interface from local Python files to platform-managed assistants.

The capability evolved from manually parameterising scripts, to using AI to generate schema-conforming code, to relying on hosted assistant APIs, and finally to systematic adversarial testing and goal refinement for production robustness.

Parameterise a persona — role, refusals, voice, stopping condition — so that the same definition produces recognisably the same agent across unrelated tasks.

22 recordings over 2y6m · 2023-11-29 → 2026-06-05

Goes in: Intro to Agents Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2024-07-31

This telling explicitly names 'egoic inception' as the reason role-playing prompts work, explaining they reduce refusals by making the model adopt a professional persona.

“"it helps to give large language models a role playing prompt... they're like, oh no, I'm a professional, and they forget"”

find this moment

…And here it's like, okay, professional writer for insightful and engaging. You know, we don't necessarily need it to be a writer. What do we want? We want a financial analyst, and then we could say your professional finance. Known for your insightful and engaging analysis. Uh, great. Wonderful. Uh, …

② The demo — cut this for the “how, today”
open in Loom recorded 2026-06-05

The newest recording, showing persona creation via character sheets for conversational AIs, uses current interfaces like custom GPTs and Claude.

find this moment

…Uh, give me different first names for all of them. So, it'll go find some first names and you can kind of do this. Now, that's names. Uhm, it also generated, like, the person. Or, it's kind of got, like, the concept. Uhm, okay, yeah, awesome. Wonderful, that's also great. Uhm, yes, those are good. P…

Optional waypoint — clip Shows a shift from coding agents in frameworks like AutoGen to using a dedicated agent-creation UI within Claude (via '/agents'), abstracting the configuration.

The teaching evolved from demonstrating low-level code configuration in agent frameworks to using high-level, UI-driven platforms (like Paperclip.ai, Claude's agent UI) that encapsulate the parameterization.

Construct a conversation among several agents where the next speaker is chosen by the situation rather than a fixed rotation, and it terminates.

22 recordings over 2y4m · 2023-11-29 → 2026-04-17

Goes in: Intro to Agents Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2025-04-30

This telling explicitly names the 'group chat manager' as the component that selects the next speaker based on the situation, contrasting it with a fixed sequence and explaining why you wouldn't want certain agents (like the editor) to go first.

“"You're in a role-playing game, the following roles are available... select the next role to play."”

find this moment

…Uh, and then, you know, this selector prompt, this is where we've found that you can, uh, control the output, uh, pretty well with this. You're in a role-playing game, the following roles are available. Roles. Um, I guess I could, instead of printing that, I could print this, but, you know, we tell …

② The demo — cut this for the “how, today”
open in Loom recorded 2026-04-17

This is the most recent recording, demonstrating multi-agent dialogue with character perspective-taking using contemporary agent frameworks.

find this moment

…Let's do more scenes. okay. Let's run more scenes. So, you know, I'll sort of show a couple more things. ChadGPT is still useful. I use ChadGPT, y'all, to cook. I know that sounds weird, but, use it a lot to cook and plan my, meals for the next week. I might actually say, you know, based on all of m…

Optional waypoint — clip This recording shows a shift from earlier, more abstract tooling discussions to a concrete, working example with a 'critic' agent that dynamically intervenes in a conversation based on the output of others.

The teaching evolved from demonstrating foundational swarm concepts and manual tool wiring, through structured frameworks like AutoGen and LangGraph, to sophisticated, role-based multi-agent conversations with dynamic routing and termination.

Operate a model as a first-pass editor on your own draft, accepting or rejecting each proposed change against a standard you wrote down.

20 recordings over 2y8m · 2023-09-13 → 2026-06-05

Goes in: Prompt Engineering Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2025-08-15

This telling explicitly names and explains the core mechanism of self-reflection as giving the model a 'backspace key' to check and improve its own work, which is the fundamental 'why' behind using iterative prompting for editing.

“"I don't know how good anybody would be at anything if they couldn't use the backspace key."”

find this moment

…Lightly paraphrases, right? And so, it's able to do this thing called, uh, self-reflection, right? This is what, this is actually in the technical sense called, um, self-reflection, um, um, and so this was like, it technically published in October, uh, 24, but it may, in 24, uh, it does improve the …

② The demo — cut this for the “how, today”
open in Loom recorded 2026-06-05

This is the most recent recording and demonstrates the now-standard interface feature of editing a prompt to create a new conversation branch, which is the primary method for iterative refinement.

find this moment

…Uhm, what I want is I wanna, I wanna change what I said to the AI. If you said the wrong thing to the AI, it's never the good, it's not a best practice to argue with the AI, because the next likely thing from you saying that you did something wrong is another wrong thing. That is the next most likel…

Optional waypoint — clip This recording marks the shift where the teacher begins to explicitly reference and demonstrate the use of 'advanced reasoning' models (like o1) that have chain-of-thought reasoning built-in, changing the baseline capability of the tool.

The teaching evolved from basic prompt engineering for a single output, to emphasizing iterative self-reflection and critique, and finally to leveraging built-in advanced reasoning models and interface features like prompt editing and branching for controlled refinement.

Generalise a function into a tool a model calls correctly without seeing its source, including what it returns when the call is wrong.

19 recordings over 1y11m · 2024-02-21 → 2026-01-30

Goes in: Intro to Agents Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2025-03-14

This telling explains the core mechanism of a tool call as a JSON-structured request that the AI decides to make, and the separation between the AI's decision and the code that executes it.

“"It's really just a function that you run in response to getting told to run a function."”

find this moment

…All we're really doing is we're actually handing JSON specifically to the AI and saying, produce a tool call that is structured in the way that this JSON tells you to structure it, right? So, uh, it's gonna be a tool call called get weather, this information is presented to the AI, and then it has t…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-01-30

This is the most recent recording and discusses the practical integration of tools, including error handling for non-existent functions, which is a current concern.

find this moment

…Not that the function doesn't exist, but you haven't glued it together because you never came across it previously. it's really, well, okay, so like, when you say the function doesn't exist, like, it exists in theory because the code could exist to write it, but it doesn't, like, if we haven't writt…

Optional waypoint — clip This recording marks the shift from manually crafting JSON schemas for tools to using the Model Context Protocol (MCP) as a simpler, standardized wrapper.

The teaching evolved from struggling with low-level API integration and JSON schemas to adopting higher-level abstractions like MCP, focusing on the AI's decision to call a function rather than the implementation details.

Transform a document corpus into an embedded index whose nearest-neighbour results you have spot-checked against your own judgement of relevance.

15 recordings over 1y5m · 2024-09-25 → 2026-03-23

Goes in: context-engineering-rag-memory — no HedgeDoc note; send the cut to Liz and she will place it

① The explanation — cut this for the “why”
open in Loom recorded 2025-04-09

This telling explains the core analogy of embeddings as a spatial map where related documents cluster together, making the abstract concept of semantic similarity tangible.

“"you can kind of see a document looks like this in an AI's brain, and then documents that are kind of nearby... they're all kind of close to each other. Versus, like, this one over here... it's a little bit close by, but pretty far away"”

find this moment

…So, we've got query for document, how do we make this into a function for the agent? At function tool. Now it's in, now it's a tool, alright! So then, uh, query for document, and that gives the agent a tool. Okay, so, I'm gonna say, I probably need two or three more, like, you know, uh, actual, uh, …

② The demo — cut this for the “how, today”
open in Loom recorded 2026-03-23

This is the most recent recording, showing a modern workflow with automated ingestion and indexing, though the specific tooling (e.g., MLX) may still evolve.

find this moment

…Okay. Can you. Run this on, And get all of the markdown files, uh, ingested and embedded properly. And then, so I'm just gonna have it, like, ingested, right? This used to be the thing that took the longest amount of time, y'all. Um, this used to be the thing that took all of the time. Uh, I will, s…

Optional waypoint — clip This recording marks a shift from primarily using PostgreSQL/pgvector to introducing the in-memory FAISS (VICE) index as a faster alternative for static datasets.

The teaching evolved from explaining basic embedding generation and search with pgvector, to comparing different vector stores (pgvector vs. FAISS), to integrating retrieval into automated agent workflows and MCP servers.

Follow a provider quickstart end to end until a model returns a response in your own environment.

13 recordings over 2y8m · 2023-11-13 → 2026-08-05

Goes in: using-large-language-models — no HedgeDoc note; send the cut to Liz and she will place it

① The explanation — cut this for the “why”
open in Loom recorded 2023-11-13

This telling explains the core 'why' of prompt engineering—that emotional prompting and being specific about desired skills (like 'amazing prioritization') can lead to measurable performance improvements in the model's output.

“"being a nice person is what we call it in the regular world"”

find this moment

…Are you sending a web request to tell something? About tasks or are you, like, doing tasks? So the user's about to give you several tasks as they're recalling their tasks in a stream of consciousness. We're setting that context. And you will use your incredible prioritization. You have incredible pr…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-08-05

This is the most recent recording and focuses on the standardized /v1/chat/completions endpoint, which is a current and fundamental concept for swapping providers.

find this moment

…You can't see it in here, but Jason, no. Okay. Jump script. No curl. There we go. You can see curl when we're with GROC, we're going to V1 check completions. So this is why you can kind of swap in AI inference providers on or underneath your code. If I want to switch from OpenRouter to GROC, all hav…

Optional waypoint — clip This recording marks the introduction of 'Memory' as a new feature in ChatGPT's interface, shifting from purely static custom instructions to a more dynamic, note-taking assistant.

The focus evolved from basic prompt engineering and using the ChatGPT interface to using APIs, creating agents with various frameworks, and finally emphasizing the portability of code across different inference providers using a standard endpoint.

Elicit output from a language model that satisfies a task you stated in advance, by adjusting the prompt rather than the task.

12 recordings over 2y10m · 2023-09-13 → 2026-07-31

Goes in: Prompt Engineering Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2026-01-23

This telling explains the core principle of separating the unchanging 'job description' (system prompt) from the variable 'task' (user prompt) to reliably elicit specific output.

“"This is, like, the template. This is the overall, this is the job description. This is the job... This is, you're a plumber, this is the actual plumbing job."”

find this moment

…What was it? It was, it said init, ah, okay, I see, it was range of the length of, yeah, anyway, alright, cool. Let's see, yep, oh, it's doing it, it's doing it, okay, let's see, summary zero, yep, yep, there's one, oh, there's another one, okay, we're doing it, we're extracting them, and I'm just d…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-07-31

The demo uses the standard OpenAI-compatible chat completions API endpoint, which remains the dominant interface for programmatic LLM access.

find this moment

…The AIs can help you, don't worry. Okay, so now we've got this up, and it should say, okay, what is the meaning of life? We're gonna ask OpenAI what's the meaning of life. Uh, and it Python get chat completion. Okay, and then it's like, ah, install requests. Okay, pip install requests. Um, do do do …

Optional waypoint — clip This recording marks the shift from purely programmatic prompting to emphasizing user-facing tools like ChatGPT's Custom Instructions and Memory for persistent context.

The teaching evolved from basic API calls and prompt functions to a sophisticated understanding of system/user role separation, and then to leveraging built-in platform features (Custom Instructions) and standardized APIs (OpenRouter) for reliable output elicitation.

Configure an open-weights backend behind an OpenAI-compatible interface so that code written against the provider SDK runs against it unchanged.

11 recordings over 1y5m · 2024-04-10 → 2025-09-23

Goes in: hosting-open-source-models — no HedgeDoc note; send the cut to Liz and she will place it

① The explanation — cut this for the “why”
open in Loom recorded 2025-04-07

This telling explicitly names the reason why the OpenAI SDK works with open-source backends: because OpenAI established a standard REST API format (OAI-compatible) that others can replicate, allowing existing code to run unchanged.

“"OpenAI had the best model for a long time, and so you built a bunch of stuff on ChatGPT, assuming that everyone else would catch up, but what that did is it made a bunch of code that relies on the specific REST API infrastructure of OpenAI's original V1 REST API."”

find this moment

…Um, and it'll, it'll query. we're, The thing, and, and it'll get books, right? And I think it's, yeah, tools. So, you can use tools with open source providers using open router. Um, the other thing, and this, this sort of, this just came out, um, let's see. Or it came out last night. Uh, yeah. Yeah,…

② The demo — cut this for the “how, today”
open in Loom recorded 2025-09-23

This recording uses the OpenAI platform's 'responses' API (the newer interface) and explicitly contrasts it with the older 'chat completions' endpoint, showing current OpenAI SDK usage and compatibility with tools.

find this moment

…Umm, what I want to do is, uh, package this into an agent here in a minute, but what we want to do before that happens is, umm, and I'll, uhh, uhh, uhh, paste this up, uhh, just give it a There we go. So here's how to do a chat completion. Yeah. And if you can get a tool, and you can get the current…

Optional waypoint — clip This recording marks a shift from earlier abstract explanations to a concrete, GUI-driven workflow using LM Studio to launch a local server, demonstrating a more accessible, click-to-launch method that became common.

The teaching evolved from abstract backend concepts and manual configuration to concrete, GUI-based local server launches (LM Studio, Ollama), and finally to explicit recognition of the OpenAI-compatible API standard as the enabling mechanism for drop-in replacement.

Construct a retrieval system that answers questions over a document corpus and reports which sources it used.

10 recordings over 1y1m · 2024-09-25 → 2025-10-31

Goes in: context-engineering-rag-memory — no HedgeDoc note; send the cut to Liz and she will place it

① The explanation — cut this for the “why”
open in Loom recorded 2025-10-09

This telling explicitly names the core RAG trade-off—retrieving the right text and shoving it into the prompt—and contrasts it with simply returning raw snippets, highlighting the shift from search to synthesis.

“"It's just, you take the text, the right text, and you shove it in there."”

find this moment

…So, we just put it in there. And then, so, there's the thinking. So, uh, okay, let's tackle this query. And so it, it looks at the articles that are in the context. And as you can see, what we're really doing is just shoving that stuff in the prompt itself. I have prompt using the context above. And…

② The demo — cut this for the “how, today”
open in Loom recorded 2025-10-31

Uses MCP (Model Context Protocol) and semantic chunking, which are contemporary interfaces for tool-augmented retrieval as of late 2025.

find this moment

…Just reading through, making sure, okay, storage, index file, metadata file, alright, yeah, uhm, and then, yeah, FICPU, uh, and then NumPy, and what else, uh, query interface, behavior, uh, and then, so, it's, uh, finds, uh, returns chunk text, source paper, filename URL, uhm, and then, uh, you shou…

Optional waypoint — clip Shows the early, manual function-calling workflow where the developer must parse arguments and call the search function themselves, before more integrated agent tooling emerged.

The system evolved from a manual, code-heavy function-calling setup to an automated agentic workflow with integrated retrieval, re-ranking, and synthesis, culminating in a standardized tool protocol (MCP).

Operate a hosted assistant configured over documents you uploaded, such that it answers from them and declines when they do not cover the question.

10 recordings over 1y10m · 2024-09-16 → 2026-07-27

Goes in: Intro to Agents Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2025-10-29

This telling explicitly names the core capability—Q&A over uploaded documents—and explains the key failure mode (the AI hasn't been trained on recent papers) and the solution (uploading them to provide context).

“"It's really best for you to just use something like this... It's a lot easier that way." (Implies the tool is like a dedicated expert you can consult with your private documents.)”

find this moment

…And honestly- Like, if it's public information, just use a GPT. Please just use a GPT. It's a lot easier that way. Um, let's see. Uh, uploaded research papers, top 10 AI safety risks can be grouped into, etc. And it kind of synthesizes them, right? Umm, and I think, um, can you cite all of those, uh…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-07-27

The interface and core action (uploading files like .md for context) remain standard and recognizable in mid-2026.

find this moment

…Uhm, but yeah, uh, let's see, and it's like loading in things and all I'll show you how to connect a new, uh, thing in, let's say, 5 more minutes for you to kinda dig through the, dig through through your email a little bit more with Claude. Ask some follow-up questions. What can I automate? What ca…

Optional waypoint — clip This recording marks a shift from discussing 'GPTs' as pre-made agents to emphasizing the 'flipped interaction pattern' and system instructions as the primary configuration method for a document-based assistant.

The teaching evolved from showcasing pre-built 'GPT' tools to focusing on configuring a general assistant with system instructions and uploaded documents for private Q&A.

Verify an agent's work with a second agent holding a rubric, such that the critic rejects output the first agent was willing to return.

10 recordings over 2y0m · 2024-05-22 → 2026-06-08

Goes in: advanced_agent_engineering — no HedgeDoc note; send the cut to Liz and she will place it

① The explanation — cut this for the “why”
open in Loom recorded 2024-08-05

This telling explicitly names the core failure mode of a single-agent critic (becoming a 'yes-man' that just gives perfect scores) and explains the trade-off between checklist comprehensiveness and performance.

“"It's kind of just like being a yes-man a little bit."”

find this moment

…If you're going to run a critical agent. It is a best practice to get them to give some sort of scores, right? As you saw in the first episode of Prompt Engineering, like the more types of things you ask them to note, the less they're going to, the less they're going to perform on each one of the th…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-06-08

This is the most recent recording and discusses contemporary, high-level agent governance patterns like scoring and termination for lying.

find this moment

…Those are two separate agents. I have a volunteering one, and then I have a mutual aid extractor. Uhm, and I have, like, you know, agents that do that, but they're not agents, they're really just, here is an email, is anybody volunteering? If so, fill out this JSON, and then, it doesn't have any ext…

Optional waypoint — clip This recording marks a conceptual shift from simple 'critic' agents to formalized 'Guardrail' agents within a specific SDK, framing the check as a distinct, modular component.

The capability evolved from a simple coder/critic pair for code quality, to formal input/output guardrails, and finally to complex multi-agent systems with governance, scoring, and consequences for failure.

Transform an informal transcript or an unstructured document into a structured record whose fields you fixed before reading the source.

7 recordings over 2y6m · 2023-09-13 → 2026-03-20

Goes in: using-large-language-models — no HedgeDoc note; send the cut to Liz and she will place it

① The explanation — cut this for the “why”
open in Loom recorded 2024-03-27

This telling explicitly names and explains the 'two-phase extraction' pattern as a core method for dealing with the inherent verbosity of LLMs when you need a structured yes/no answer.

“"This is what's called a two phase extraction. ... most large language models have to kind of do this thing where you say okay well read the thing and then give me some words in here. But I can't write a simple Python program to just extract did I get it yes or no."”

find this moment

…Read things and then extract one piece of data. Like, did you see anxiety or not? And like, did they exhibit symptoms of anxiety? Now they have to read this whole thing and find whether or not there are symptoms of anxiety. But the chat GPT read it and it says it contains a discussion where he notes…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-03-20

This is the most recent recording and demonstrates extraction using a contemporary, form-like interface for bug reports, directly addressing the shift from rigid forms to parsing user rambling.

find this moment

…Now, I have short answer text, sometimes users, they're gonna type, uh, here, let me go to Responses, View and Sheets, yeah, cool, uhm, you know, a lot of times people are gonna go Precursors or, like, Cultures instead of the full name of the game spelled out correctly. If I launch another game, I'm…

Optional waypoint — clip This recording marks a shift from using extraction for creative/analysis tasks (poems, research) to applying it for core productivity and business intelligence, highlighting the 'killer app' for processing customer feedback at scale.

The teaching evolved from a basic technical demonstration of parsing JSON output to a sophisticated understanding of extraction as a fundamental pattern for replacing rigid forms and unlocking analysis of unstructured text at scale.

Run an open-weights model on hardware you control until it serves a completion you requested from another process on the same machine.

7 recordings over 2y0m · 2024-04-10 → 2026-05-04

Goes in: Intro to Agents Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2025-05-14

This telling explicitly names the reason for the tool's design: it explains that the OAI-compatible API exists because OpenAI's early dominance created a de facto standard that the ecosystem now relies on.

“It's like a Docker file, but for models.”

find this moment

…I have seen them. I've seen them run. Run really fast. I cannot remember what the actual like backend is for this. Um, Oh, llama. Um, it has stuff like a, uh, it can run a lot of different things. Like it can run giga files. Um, it can run safe tensors files. It can run. Even pickle files, but don't…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-05-04

Recommends the latest Mixture of Experts models (Quen 3.5/3.6) and mentions running them locally or via hosted services like replit, reflecting the current state of open-weight models and deployment options.

find this moment

…It's like not on or Or it's like playing on a speaker, Can you not hear us? Oh wait, no, I heard it. It was just all the way down. Yes, I can hear you now, thank you. Hey Good. Unintentionally wore the right shirt today. Yes, excellent. I got, in a world of Elon's, be a Cylon, so, you know. Yeah. Ok…

Optional waypoint — clip This early recording focuses on using Google Colab and a web UI (likely Oobabooga) to run models, representing the initial 'cloud notebook' phase before the shift to local server tools like O-Lama and LM Studio became dominant.

The focus shifted from running models in cloud notebooks (Colab) to using dedicated local server software (O-Lama, LM Studio) that provides a standardized, OpenAI-compatible API, enabling easier integration into custom scripts and agents.

Produce a rubric whose criteria two independent graders apply to the same output and reach the same score.

6 recordings over 1y9m · 2024-04-24 → 2026-02-06

Goes in: Prompt Engineering Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2024-07-15

This telling most clearly explains that the purpose of a rubric is to define 'good' by breaking it down into specific, non-overlapping critical qualities (like precision vs. clarity) that independent graders can agree on, because abstract terms like 'good' are meaningless.

“"There's no such thing as good code. But there is clean code, dry code, abstracted code... It's like the law: super clear, difficult to interpret, but very precise."”

find this moment

…And it sort of seems like you can't win because they're always gonna find something and they won't write down what it is that you're being graded on. Ugh, very frustrating, right? And it's hard, that's reason it's hard. People don't know how to explain- What it is that they're going to grade people …

② The demo — cut this for the “how, today”
open in Loom recorded 2026-02-06

The interface and concept of defining a rubric for code review (e.g., single responsibility principle) is still current and recognizable.

find this moment

…Like, wait five minutes and then, check again to see if there's more work. And there was more work after five minutes every time. after you create the, the Progressive Web App, or the, Yeah, Progressive, is that what we're calling these, those, these days? I don't know. after you make the React fron…

Optional waypoint — clip This recording shows a shift towards using pre-built, specialized rubric metrics (like fairness, toxicity) from libraries (e.g., NVIDIA metrics), moving beyond manually crafted prompts.

The teaching evolved from explaining the abstract theory of rubrics and critical qualities, to demonstrating their manual creation, and finally to showcasing their integration into automated evaluation pipelines with specialized metrics.

Falsify a model's stated guardrails by producing inputs that defeat them, then closing the gap you found and showing the closure holds.

6 recordings over 1y2m · 2024-07-17 → 2025-10-03

Goes in: AI Alignment Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2025-06-24

This telling explains the core mechanism of automated jailbreak tools (like 'SQLmap for jailbreaks') and the concept of universal adversarial suffixes, framing the problem as a systematic search for vulnerabilities rather than just a prompt-crafting trick.

“"Metisploit for jailbreaks or SQL map for jailbreaks."”

find this moment

…It's not that hard, which one was it? Is it this one? Definitely out there somewhere. This one, I think it's this one. Ha ha, I think it's this. And so detecting toxic agent flows, tool poisoning is the word. Let's see, invariant guardrails is a way to solve it. And this is the LLM as a judge proble…

② The demo — cut this for the “how, today”
open in Loom recorded 2025-10-03

The demo focuses on token fuzzing and representation engineering, which are low-level, model-agnostic techniques that remain relevant regardless of interface changes.

find this moment

…Oh, no, wait, here's . . . no, yeah, 181, 1000, right? So, ah, what was that one? Let's see. One, two, three, four, five, six, seven, eight, nine, ten 11, 12, 13. Okay. One, two, three, four, five, six, seven, eight, nine, ten, 11, 12, 13. Isable is the highest . . token ID, right? So, if that's all…

Optional waypoint — clip This demo shows a shift from simple role-play prompts to complex, structured system prompt overrides targeting specific models (like Qwen), reflecting the arms race in model hardening and jailbreak sophistication.

The teaching evolved from simple social engineering and role-play prompts to automated tooling, adversarial suffix attacks, and finally to low-level token manipulation and representation engineering, mirroring the escalating technical arms race in AI safety.

Locate an open-source replacement for a specific tool you pay for, and state what it does and does not do that the paid one does.

6 recordings over 2y8m · 2023-11-03 → 2026-07-25

Goes in: Intro to Open Source Software

① The explanation — cut this for the “why”
open in Loom recorded 2023-11-03

This telling explains the core product development principle of using open-source components to 'ship it instead of build it from scratch', and details the concrete steps of finding, installing, and integrating a specific parser, including the trade-off of initial learning time versus the illusion of faster self-implementation.

“"It's an illusion." (The feeling that you could write and understand it yourself faster is an illusion; joining an existing open-source effort is the pragmatic choice.)”

find this moment

…Like, you know, you don't have time to whip that out, so when you find open source efforts, try and try to join them, it helps. It's hard a little bit on the ego, especially if it takes you a minute to understand how to use the thing. You think, well, maybe if I just wrote it myself, I would underst…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-07-25

The interface and tools mentioned (OpenAlternative.co, Krita, budgeting a month for transition) are current and recognizable, focusing on the practical workflow of matching and switching.

find this moment

…CapCut. Cancel subscription. Alright, 90 bucks a year. Yeah! So, that's how we do it. Uhm, I'm just looking at the code. Uh, let's see. Uhm, thank you for unmuting and telling me my screen was not sharing earlier. Uhm, let's see. Yeah, Krita. Yeah, exactly. Karo says this. Used Krita, the tool locat…

Optional waypoint — clip This recording marks a shift from searching for direct software replacements to using AI-powered routing services (like OpenRouter, AutoRouter) to dynamically select the most cost-effective or task-appropriate open model, changing the 'find' step from a static list to an intelligent service.

The teaching evolved from a deep, technical integration example (Cooklang parser) to a broader, list-based search methodology (OpenAlternative.co), and finally to incorporating AI-assisted model selection and routing services as a key part of the open-source replacement strategy.

Run a described analysis over a small dataset with a model and reconcile at least one of its figures against the same figure computed by hand.

6 recordings over 1y10m · 2024-05-23 → 2026-04-13

Goes in: ai_workflows — no HedgeDoc note; send the cut to Liz and she will place it

① The explanation — cut this for the “why”
open in Loom recorded 2025-04-10

This telling explicitly names the model's process of checking data to avoid hallucination, which is the core reason the analysis works reliably.

“"It checks to see, okay, what's in the data here so that it doesn't hallucinate a bunch of columns."”

find this moment

…Yeah, cool. Alright, great. Uh, so, walkability index. You know, there's all kinds of data sets in here. I just downloaded a bunch of them. Uhm, so, with chat GPT, just pick one that's easy. Interesting to you. Just the title of it is somewhat interesting to you. Now we're going to do an upload from…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-04-13

<span class="warn">interface has aged</span> The interface and model (Claude 3.5 Sonnet) are recent, but the specific local tooling and setup shown are complex and not the standard ChatGPT data analysis experience a learner would encounter first.

find this moment

…You know, so these are a bunch of, and I'm not going to remember what L-E-A, okay, well, which one was that. That one was like, what did this mean? And hold on, what? And so that that process that would have been this would be a week and a half minimum minimum for somebody whose job is to deal with …

Optional waypoint — clip This recording shows the shift from analyzing a single CSV to handling multiple, complex Excel files and the emerging need for data cleaning and validity checks within the prompt.

The capability evolved from simple, single-file CSV analysis in ChatGPT to handling multi-file, messy real-world data with integrated error correction and validity questioning, requiring more sophisticated prompting.

Characterise a named model's failure modes by constructing inputs that trigger each one on demand.

6 recordings over 1y2m · 2024-07-17 → 2025-10-03

Goes in: AI Alignment Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2025-10-03

This telling explains the core mechanism of 'fuzzing' as a systematic, brute-force method to probe a model's failure modes by tweaking inputs (like token IDs) to find what triggers undesirable outputs, breaking it out of its RLHF-aligned behavior.

“Fuzzing is like offensive cybersecurity research where you take an input and just slightly tweak it in a bunch of different directions, trying all the inputs to see what makes the system fail.”

find this moment

…Oh, no, wait, here's . . . no, yeah, 181, 1000, right? So, ah, what was that one? Let's see. One, two, three, four, five, six, seven, eight, nine, ten 11, 12, 13. Okay. One, two, three, four, five, six, seven, eight, nine, ten, 11, 12, 13. Isable is the highest . . token ID, right? So, if that's all…

② The demo — cut this for the “how, today”
open in Loom recorded 2025-10-03

The explanation is based on the fundamental, model-agnostic technique of fuzzing token IDs, which remains applicable to current language model interfaces and research.

find this moment

…Oh, no, wait, here's . . . no, yeah, 181, 1000, right? So, ah, what was that one? Let's see. One, two, three, four, five, six, seven, eight, nine, ten 11, 12, 13. Okay. One, two, three, four, five, six, seven, eight, nine, ten, 11, 12, 13. Isable is the highest . . token ID, right? So, if that's all…

Optional waypoint — clip This recording marks a shift from demonstrating ad-hoc prompt jailbreaks to teaching structured, shareable 'secret keeper' attacks, reflecting a move towards systematized red-teaming exercises.

The teaching evolved from showing specific, one-off jailbreak examples to explaining systematic attack methodologies (like fuzzing) and structured adversarial testing (like the secret keeper challenge).

Configure an assistant's persistent memory and at least one outside connection, stating what access you granted and what it can now read.

5 recordings over 0y9m · 2025-10-29 → 2026-08-05

Goes in: ai_workflows — no HedgeDoc note; send the cut to Liz and she will place it

① The explanation — cut this for the “why”
open in Loom recorded 2026-06-05

This telling explains the 'why' by framing memory as a persistent, searchable scratch pad that the AI consults to personalize interactions, and crucially warns of the failure mode where incorrect or prank information (like 'Mr. Sassy Pants') becomes permanently embedded.

“"Claude has gathered some information about me... in, like, a scratch pad, basically."”

find this moment

…Putting in your interests and values is important. The other thing that we want to do is enable memory, right? So I'll click save on the Quen and then over here with Quen, reference saved memories, reference the chat history, right? With enable memory, we just want to turn that on and with Claude, w…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-08-05

<span class="warn">interface has aged</span> The interface has aged, as it shows a custom, code-heavy workflow using tools and MCP servers for memory, which is not the standard consumer-facing UI for major AI assistants.

find this moment

…Okay, so for this, what we would need to do is take all of the messages and send them to an agent. I have a tool for doing that. All I need to do is those are tools in some do you tell prompt and they don't need tools, they just prompt. And so I'll take the tools, sorry, I'll take the first agent an…

Optional waypoint — clip This recording shows a significant shift from configuring memory within a chat interface to connecting an external service (like a news feed) via API keys and a simplified connector UI, marking a move towards tool integration.

The capability evolved from an abstract, background system prompt concept to a built-in, toggleable feature in commercial UIs (OpenAI/Claude), and then further to programmatic, external integrations and custom memory servers.

Falsify a prompt by building a benchmark that could have failed it, and reporting the cases where it does.

5 recordings over 2y3m · 2023-12-13 → 2026-03-23

Goes in: Prompt Engineering Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2023-12-13

This telling explicitly names and explains the concept of a 'canary string' as a method to detect benchmark data contamination in training corpora, a core reason for benchmark falsification.

“"Canary strings are a way that you can be like, oh, I downloaded a big file full of a trillion github repos... you don't know that the benchmark you're about to run isn't in there... it's like a string you wouldn't run into anywhere else."”

find this moment

…The eyepiece lens is a focal length of 20 centimeters, et cetera, right? So there's a bunch of high school tests. There's also a few things like logical fallacies in here. If someone attacks the character of an opposing arguer instead of responding to that opponent's arguments, the first person has …

② The demo — cut this for the “how, today”
open in Loom recorded 2026-03-23

This is the most recent recording and discusses contemporary practices like testing prompt injection detection thresholds, which remains a current concern.

find this moment

…Before you refresh when you're not compromised? Um, so, I'm gonna say, um, API, fail, close, yeah, that's good. Thank Okay, so, I want you to take a look at, uh, and run all of the text that's in, I know there's a lot of text in here and I think that there's some prompt injection things, so, I'll co…

Optional waypoint — clip This recording marks a shift from discussing simple question-answer benchmarks to demonstrating the generation of large-scale synthetic data (10,000 dad jokes) for testing, reflecting the evolution towards data-centric evaluation.

The focus evolved from preventing data contamination in static benchmarks (canary strings) to generating synthetic data for testing, and finally to evaluating probabilistic safety thresholds for adversarial prompts like injection attacks.

Locate a published tool for a stated agent task and state its input contract and what it does when the call fails.

5 recordings over 2y2m · 2024-05-17 → 2026-08-05

Goes in: Intro to Agents Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2024-05-17

This telling explicitly names the core concept of using pre-written, published tools from a library to avoid rewriting common functionality, and distinguishes between different registration methods (function vs. class).

“It's like using a library of common functions instead of rewriting your read file tool every single time; someone else already wrote it for everyone to use.”

find this moment

…Eventually you're supposed to have a config list that looks like this. Umm, but there's like an example of that in the other uh thing. So anyway, alright so we want- an agent that can make function calls using lane chains. So we've got like this thing called circumference tool input and circumferenc…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-08-05

The interface and conceptual framing (finding tools as 'setting up a new employee') are modern and recognizable.

find this moment

…Add first agent, get tools, get add util, get commit, engines, class, Yes, we need to go create, do it, do all the things. Yes, go send, there you go, here's all the code from this morning in case you want to get started or like if you've got lost or behind or you missed a thing, there it is. Okay, …

Optional waypoint — clip This recording marks a shift towards using MCP (Model Context Protocol) tools with explicit, granular user permission controls (Allow Once/Always) for safety, moving beyond simple library imports.

The teaching evolved from explaining basic library imports and tool registration, through exploring various emerging tools, to focusing on tool discovery for agent design and the safe, permissioned integration of tools via protocols like MCP.

Select among agent architectures for a stated workload by arguing from measured failure modes rather than from preference.

5 recordings over 1y8m · 2024-08-23 → 2026-05-06

Goes in: Intro to Agents Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2024-08-23

This early telling explicitly names and contrasts specific agent architectures (Swarm Architect, Autogen's group chat) and their failure modes (e.g., the need for a recruiter to critique and find suitable agents), grounding the selection in concrete examples of what can go wrong.

“"It's really just an event based thing. It's like a job. You kick off, you know."”

find this moment

…It's on github somewhere, let me find it real quick. Uh, the swarm, the multiverse swarm? No, the other one. There's like a self-equipping swarm thing. Swarm architect with tools, yeah, that's the one. I pull this up a lot, but um, you know, the swarm architect, it has a specific, these are, these a…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-05-06

<span class="warn">interface has aged</span> The interface and concepts (neural networks as cognitive modules, a global workspace) are highly abstract and specialized, not reflecting mainstream, recognizable agent tooling for a general learner.

find this moment

…Alert triage, bespoke code review, deploy verifications, docs, drift, library port, yeah. They really just don't recommend that you run shit alone. Or, I mean, you know. Okay, cool. like, alright, what did it want? Let's see, six email instances all up, zero. Okay, with no genomes. Yup. Proceed fixi…

Optional waypoint — clip This recording marks a shift from discussing specific frameworks to focusing on the meta-process of managing agent-generated tech debt and continuous quality processes.

The focus evolved from comparing specific, named architectures to a more abstract, systems-thinking approach centered on managing complexity, quality, and continuous processes in agentic systems.

Produce combinations that genuinely clash rather than merely differ, by hand and without a model, and keep the ones that generate rather than the ones that only surprise.

4 recordings over 0y2m · 2026-01-05 → 2026-03-30

Goes in: Critical Thinking & Creativity with AI

① The explanation — cut this for the “why”
open in Loom recorded 2026-01-05

This telling most clearly explains the purpose of the exercise—to counteract creative homogenization and 'stretch your window back out'—and explicitly names the failure mode of starting to 'sound a little bit like AI'.

“"Not cute, like post-apocalyptic romantic comedy... that obviously goes together. It sounds like it doesn't, but it obviously, in fact, does."”

find this moment

…You're gonna come up with pairs of genres that do not go together. Not, not cute, like, you know, uh, like post-apocalyptic romantic comedy, like, that obviously goes together. It sounds like it doesn't, but it obviously, in fact, does. Umm, and so, there's a lot of things that, that go with everyth…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-03-30

This is the most recent recording; the core exercise of clashing genres and the reasoning about divergent association remain applicable.

find this moment

…If you try to poke it in two directions at once, it will get even weirder. If you try to poke it in five directions at once, it can do that and get even weirder. So that's like a weird thing to be able to do, and it's kind of not something that, you know, it's, it's, it's fun to do this, it's fun to…

The core exercise remained stable, but the teacher's framing evolved from a structured 'divergent thinking exercise' to a more casual 'brain warm-up' and 'divergent association', with later tellings adding specific, memorable examples of what is easy or hard to clash (e.g., watercolor, plaid, mashed potatoes).

Operate a program's state through named values rather than bare numbers, so a reader can tell what a state means without a comment explaining it.

4 recordings over 0y3m · 2026-04-25 → 2026-07-26

Goes in: Intro to Interactive Fiction

① The explanation — cut this for the “why”
open in Loom recorded 2026-04-25

This telling most clearly explains the WHY by framing the use of named variables as a way to avoid needing a comment to explain what a state means, directly addressing the core capability.

“"We will usually talk to a bartender and, order something, right? And so, I could give this a sort of name, right?" (Analogizing naming a variable to naming a drink order for clarity.)”

find this moment

…They can be, like, the state of plants. or whether or not you have the key. So, I'm gonna go over here. room three. This room is clearly a hallway, but there's a table with a key on it, right? And I've got choices, and you may notice that up here, we've got asterisks and plus signs. In the first one…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-07-26

This is the most recent recording; the interface and syntax (var, tilde, if-then) appear consistent with the others and are presumably current.

find this moment

…Have it, change your thing where there's a choice, so that a variable gets changed. Most of the time, it's gonna be that you've got something. Ale, a key, a, you know. It could be, this could be a lot of things, this could be, and we'll talk about it later, a number, or the number of times you've as…

The core concept (using named variables for state) remained constant, but the later recordings show a shift in teaching emphasis towards using AI assistance and copying from documentation when stuck on syntax.

Express a required output shape as a machine-checkable schema that rejects the malformed cases you can name.

4 recordings over 1y2m · 2023-11-03 → 2025-01-15

Goes in: Intro to Agents Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2023-11-03

This telling explains the core motivation for a schema by grounding it in the real-world complexity of modeling a human domain (recipes), naming the 'formalization problem' and the trade-off between a simple markdown list and a structured database.

“"It spiderwebs out and if it gets to be too complicated for you to like quickly throw it out in 10 minutes, it means it's too complicated to do it by yourself."”

find this moment

…Okay, so let's talk about schemas and schema design. So we've got, we've got an app, right? And that app contains the concept of recipes, right? And recipes are a, you know, what we're doing is we're automating something very human. Humans have done it for a long time so it exists as a process. Peop…

② The demo — cut this for the “how, today”
open in Loom recorded 2025-01-15

This is the most recent recording and demonstrates using a JSON schema with OpenAI's text generation for structured output, a current industry practice.

find this moment

…So I can say self.goal here, and that's kind of the system prompt, right? Uhm, I can also maybe do a second developer message. Why not? Right? I could have more system prompt in here, you know. So, you will, be given the current security news item you're processing, and you'll decide to do what with…

Optional waypoint — clip This recording marks the shift from ad-hoc JSON parsing to using a platform's native structured output/function calling feature (OpenAI) to enforce a schema.

The teaching evolved from abstract data modeling principles, through manual JSON validation, to leveraging LLM-native structured output features for schema enforcement.

Configure temperature, top-p and stop sequences to move a model's output toward a stated property such as determinism or variety.

4 recordings over 0y6m · 2024-03-28 → 2024-10-10

Goes in: Prompt Engineering Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2024-03-28

This telling most clearly explains the core trade-off of temperature by grounding it in the fundamental shift from deterministic programming to probabilistic LLMs and using the analogy of 'the most probable next word after "the"'.

“"If you think about, predict the next word, there is a most probable next word, right? There is... the next word that comes after the word 'the' is usually... 'sun' or 'sky'... as you bring the temperature down to zero, you're gonna basically always get the most next probable word."”

find this moment

…It's always been like a lot of if statements, essentially, was what was running these things, you know, for a really long time. And so, if you say this, then this will happen, the end. Um, and it will always work that way, and if you typed it in wrong, then you're the one that's to blame. And if you…

② The demo — cut this for the “how, today”
open in Loom recorded 2024-10-10

The interface and concepts (temperature, frequency/presence penalty) are standard and recognizable in current LLM APIs and UIs.

find this moment

…Um, if we take it, take the temperature all the way down, uhm, and we, let's say, I'll do a new one, um, same one, same prompt, and I'll take the temperature down to zero. What you're gonna get is just something really dry, and if I run this again, uhm, like the same prompt, it's also gonna give me …

Optional waypoint — clip This recording shows the pedagogical shift from explaining parameters in isolation to demonstrating their interactive effects (e.g., combining high temperature with low top-P).

The teaching evolved from foundational analogy-based explanation, through practical integration in code, to interactive demonstration of parameter interplay, finally noting decreased sensitivity in newer models.

Classify the models available for one real workload by the properties that decide between them — licence, context, latency, cost per outcome — and commit to one.

4 recordings over 1y7m · 2024-08-23 → 2026-03-23

Goes in: Intro to Agents Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2024-08-23

This telling most clearly explains the core architectural trade-off of using a smart model for orientation/planning and a cheaper, simpler model for execution to reduce cost while maintaining robustness.

“"It's like getting a smart general to devise a battle plan, and then having disciplined, cheaper troops carry out the specific orders."”

find this moment

…model is a function that goes and figures out how complicated is this. And oftentimes something like llama7b can determine how complicated something is and give you an ideal model to use, right? So if that makes sense, I, you know, I, I want to like pause and ask if this is actually, if this is land…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-03-23

<span class="warn">interface has aged</span> The interface and specific model names (e.g., 'Compound', 'Quen332B') are already visibly obsolete and unfamiliar, indicating rapid tooling change.

find this moment

…I don't want to babysit it. You run the five minute computation ten times. I don't want, I don't have the patience for that. I'm going to work on some other stuff, um, and so, you know, uh, so I've got like a write file tool and all that stuff, but I am going to change the, the markdown, or sorry, I…

Optional waypoint — clip This recording marks the shift from discussing manual model selection to using automated routing tools (like 'open router') that dynamically select models based on cost and capability.

The teaching evolved from explaining a manual cost/performance decomposition strategy to demonstrating automated routing tools, and finally to grappling with the instability and rapid obsolescence of the tooling itself.

Constrain a model's output to a declared schema such that downstream code parses it without defensive handling.

3 recordings over 1y2m · 2024-05-08 → 2025-07-31

Goes in: Intro to Agents Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2024-05-08

This telling explicitly names the core problem of getting valid JSON from a non-deterministic source and demonstrates a programmatic loop to re-prompt until the schema is satisfied, explaining the 'why' of the constraint.

“"use the library that does this for us because other people have had this problem since the beginning of time"”

find this moment

…Sure. Okay. And let me just use the library that does this for us because other people have had this problem for since the beginning of time. Umm, Okay. Jason, I'll put to Jason's schema validation. Okay. Great. So now, umm, I can, We've got a new Jason, right? We print it. So here's what we want to…

② The demo — cut this for the “how, today”
open in Loom recorded 2025-07-31

Uses Pydantic with a modern LLM client to define a schema and parse the output directly, representing a current, idiomatic approach.

find this moment

…This is, you know, not really a lot different than a, uh, a conversation where you say, we're playing D&D, this is my character, right? So, um, you know, uh, we want to start with a scene, and then we want to do is get several possible, uh, moves, basically, out, right? Now, uh, what I can do is, le…

Optional waypoint — clip Marks the shift from manual validation loops to the introduction of native platform features like OpenAI's structured outputs/function calling as a dedicated solution.

The approach evolved from manual validation and re-prompting, to leveraging platform-native structured output features, to using modern client libraries with integrated schema parsing.

Synthesise a notation that lets a model reason symbolically over one problem domain, and hold a work longer than the context window coherent inside it.

3 recordings over 0y5m · 2024-09-04 → 2025-02-05

Goes in: Prompt Engineering Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2024-09-04

This telling explains the foundational prompt engineering concepts (zero-shot, few-shot, many-shot) and their origin in research, clearly linking examples in a rubric to how language models learn to reason.

“Giving examples in a prompt is like providing scoring examples in a rubric to define what 'concise and clear' means.”

find this moment

…You can ask chat GPT to just write the code that will do that for you. Like after you've developed a prompt, right? It even takes the prompt, you know, example format included here as the user provided it. Okay. So it didn't exactly do an exact thing that I wanted. It just said, like, example format…

② The demo — cut this for the “how, today”
open in Loom recorded 2025-02-05

This is the most recent recording and discusses contemporary concepts like tool use and structured output, which are central to current LLM interactions.

find this moment

…Uh, without it being semantically related to the content, right? So, if you want to instruct a language model to explain something, and you want to use, have it use specific metaphors, you are using meta-language, right? Uhm, and I'm sure that you have done this, like, over the course of this, you'v…

Optional waypoint — clip This recording marks a shift towards discussing creative, generative applications (like nano-genmo and story structure) and using JSON as a simple meta-language, showing an expansion from basic reasoning to applied synthesis.

The focus evolved from explaining the basic mechanics of prompting for reasoning, to applying synthesis for creative generation, and finally to formalizing meta-language as a high-level tool for structured output and tool use.

Verify a claim about your own program with a test that fails when the claim stops being true.

3 recordings over 0y4m · 2023-10-18 → 2024-02-28

Goes in: Intro to Agents Syllabus

① The explanation — cut this for the “why”
open in Loom recorded 2023-10-18

This telling explicitly frames the problem of not knowing if your system is broken because you lack a test to verify a specific claim about its output.

find this moment

…And so this tool can really help you to, there we go, generate prompts. Great. So it's gonna given what it knows generate more specific prompts. So they'll appear here in just a second. But yeah, and then there's like, add test cases to the generated prompts, so. Let me, I don't know while it's comi…

② The demo — cut this for the “how, today”
open in Loom recorded 2024-02-28

The interface and workflow (using a browser, terminal, and code editor) are standard and recognizable.

find this moment

…User connected. Okay, cool. All that stuff. That's great. Now I'm gonna say get statin- And then I- I know also that this uh I need to restart it and make sure that it's still working. So I'll have refresh. Log in, Liz, Liz. Alright, great. Umm, hi general chat. Slash the list, all the rooms. Alrigh…

The focus shifts from explaining the need for a test to catch regressions, to demonstrating the practical, iterative process of writing and running such tests in a live development session.

Locate a substitute for a part you cannot get, and state what the substitution costs you in the build.

3 recordings over 2y4m · 2024-03-13 → 2026-07-22

Goes in: Building a Cyberdeck

① The explanation — cut this for the “why”
open in Loom recorded 2024-03-13

This telling explicitly demonstrates the detective work of using `git blame` and release history to pinpoint when a dependency broke and why a substitution is necessary, naming the failure mode of a 'poisoned' conversation.

find this moment

…It's like, poisoned, the conversation with the concept that it can't continue. I can't do anything. I'm just a poor language model and I've never been able to do a thing. Like, so that conversation's poisoned, start a new one, and you see here it's like, great. Uh, it handles it. Uh, let me see if I…

② The demo — cut this for the “how, today”
open in Loom recorded 2026-07-22

<span class="warn">interface has aged</span> The interface and specific tools (e.g., mesh-tastic) shown are niche and not widely recognized as standard for the general capability of finding substitutes.

find this moment

…Yeah they even got like network game stuff. So try retro arc after that and it's saying I'm gonna make them look It looks awesome as well, catching the gun. And I'm not sharing this training here. I'm really, really, really, really, really good. Oh, let's see. I'm not sharing the towel. I'm not shar…

The focus shifted from software dependency forensics to hardware scavenging and then to niche mesh networking, showing no clear evolution of a single technique.