Today, we are covering the #1 most requested topic among Neuron readers. Seriously, this is the number one question we get every week:
How do I actually use AI agents to save time in my business?
To answer this question, we went live with James McAulay, founder of the Agentic Growth Accelerator, to assemble those pieces in public.
James built his company around only himself and Claude Code, has trained more than 400 people to date sine going solo, and publishes hands-on agent tutorials on his YouTube channel (which you should totally check out).
In the interview, James takes us on a practical crash course on everything you need to build helpful, proactive agents in Claude Cowork and Claude Code. You can watch it in full below.
This companion guide reorganizes the full two-hour stream into a sequence you can follow. Every major section links to the exact moment in the video. Watch the whole thing, or jump straight to the part holding up your workflow.
As one viewer said, the goal with agents is basically to build Tony Stark’s J.A.R.V.I.S. That is still ambitious, at least for today's AI tech. The version of agents we showed off is closer to an extremely capable chief of staff ...who occasionally needs you to reconnect Notion.
Now, this guide assumes you already understand ordinary AI chat. Readers starting from zero should begin with our AI for Total Beginners companion guide or the broader five-level AI proficiency framework.
Now, let's get into it.
So, most people asking about AI agents are trying to skip directly to the exciting part: a digital employee that wakes up, checks the business, completes useful work, and sends back the finished result.
That is also where most agent projects fall apart.
James says a useful agent requires four layers: context, connections, repeatable skills, and a reliable trigger. Skip the context, and it behaves like an intern who missed onboarding. Skip the tools, and it can only offer advice. Skip the skill, and every run becomes an improvisation. Skip the trigger, and the “proactive” agent still waits for you to remember it exists.
We're going to dive into all of that as we go, but here's a simple framework you can internalize and take with you as you read through the rest of the guide.
- Guide map
- First up, the TL;DR
- Before you start: Pick one workflow and one tool
- Step 1: Stop thinking of an agent as a magical chatbot
- Step 2: Give the agent a second brain before giving it a job
- Step 3: Connect the tools you already use
- Step 4: Design permissions before chasing autonomy
- Step 5: Turn repeated work into skills
- Step 6: Test the skill before putting it on a schedule
- Step 7: Put a reliable skill on a schedule
- Step 8: Move from reporting to proactive work
- Step 9: Choose the right model for each part of the job
- Step 10: Scale the knowledge system, not the context window
- What this guide cannot automate away
- Advanced path: Deploy the agent as a cloud service
- Your action plan for the next seven days
- Full key moments from the livestream
- BONUS Q&A: Questions from the live chat
- All resources and links
Guide map
- Pick one narrow workflow and one tool.
- Define the finished outcome an agent should deliver.
- Build a verified second brain.
- Connect the minimum tools the job requires.
- Remove dangerous permissions before adding autonomy.
- Package repeated work into a skill.
- Test the quality floor with historical and simulated examples.
- Schedule the reliable version.
- Upgrade reports into recommendations and drafts.
- Match model cost to each part of the job, then scale the knowledge system.
The advanced section down below covers cloud deployment, followed by a seven-day action plan, audience Q&A, and every resource from the stream.
First up, the TL;DR
Here are the most useful moments from the stream:
- What an AI agent actually is (5:19): James defines an agent as an AI model, context about you or your business, and tools that let it act.
- The shift from chatting to delegating (5:52): The goal is a system that returns with the report finished, the email drafted, or the task completed.
- Why context changes everything (8:10): A brilliant CFO with no company data can only give generic advice.
- The beginner-friendly starting point (12:14): James recommends Claude Cowork as the bridge between ordinary chat and Claude Code.
- Build your second brain (14:41): Create a compact folder containing your goals, role, company context, work style, and instructions.
- What Claude.md does (19:51): This file tells the agent who it is, what the workspace contains, and how it should behave.
- MCP explained in plain English (28:12): Connectors translate your request into an action inside another app.
- The Todoist demo (30:33): Claude creates and edits a real task from a normal sentence.
- A practical starter workflow (34:40): Connect meeting transcripts to task management, then let the agent extract commitments.
- The first security rule (39:33): Remove destructive permissions instead of merely asking the model to behave.
- How to organize a huge knowledge base (44:14): Use indexes, summaries, and links so the agent can find the right file without reading everything.
- What an agent skill is (46:39): A skill packages a repeatable workflow into instructions the agent can reuse.
- How to test a skill before trusting it (58:48): Run simulated users or historical examples through it, then inspect where it fails.
- Cloud versus local scheduled tasks (1:01:45): Cloud tasks run while your laptop is off, but they cannot reach files stored only on that laptop.
- The four levels of proactive agents (1:12:05): Reporting, recommending, drafting, and self-improving.
- Why a private agent inbox is safer (1:19:57): Connecting a powerful agent to your public inbox creates a prompt-injection surface.
- Which model should do which job (1:27:30): Use the strongest model for planning and skill design, then cheaper models for routine execution.
- A cloud-hosted health agent built in 30 minutes (1:53:34): James uses Vercel Eve to turn a folder into an agent that messages him through Telegram.
We share the full insights from the whole video towards the end of this guide.
Before you start: Pick one workflow and one tool
James recommends Claude Cowork for beginners because it sits between ordinary chat and the more technical Claude Code. The framework still transfers to ChatGPT, Copilot, Gemini, and independent agent tools: persistent context, connected tools, repeatable instructions, permissions, and a trigger.
Start with one job that has a clear input and a visible finished result. Good examples include turning meeting transcripts into tasks, preparing a weekly report, researching leads, or building a morning schedule. Avoid payments, deletions, public posting, and unrestricted email until the workflow has earned trust.
Your subscription should match the workload. James started at $20 and moved upward as his usage grew. Build the manual version first. Upgrade when rate limits block useful work, not because an expensive plan makes the agent sound more official.
⚡ Action step: Write one sentence in this format: “Every [day / week], take [specific input] and produce [specific output] for me to review.”
Step 1: Stop thinking of an agent as a magical chatbot
James’s definition is the cleanest starting point:
An agent has a brain, context, and the ability to do things.
The “brain” is the large language model, such as Claude, ChatGPT, Gemini, or another model. The context is the information it needs about you, your company, the task, and the rules. The tools are the connections that let it read or change something outside the chat.
A normal chatbot can explain how to update your customer relationship management system. An agent can read the meeting transcript, update the deal stage, create the follow-up task, and draft the invoice for approval.
That distinction appears early in the stream at 5:52. James wants people to move from talking about work to delegating work.
The simplest test is the final sentence:
- Chatbot: “You should follow up with this prospect.”
- Agent: “I drafted the follow-up, updated the CRM, and created a task for Friday.”
You still review the result. The agent earns more freedom as it proves reliable.
⚡ Action step: Rewrite one recurring request as a finished outcome. Replace “help me with my sales calls” with “review today’s sales calls, update the CRM, and draft the follow-ups for approval.”
Step 2: Give the agent a second brain before giving it a job
The best model in the world cannot infer your priorities, team structure, tone, constraints, or definition of good work from one sentence.
James explains this with a CFO analogy at 8:10. Imagine meeting a brilliant finance chief for coffee and asking what to do with your company. With no accounts or operating data, the advice will sound like a textbook: increase revenue, reduce costs, preserve cash.
Give that same person read-only access to Stripe, QuickBooks, and your forecasts before the meeting. The conversation can begin with a specific problem.
Agents work the same way.
James’s “Set Up My Second Brain” skill interviews the user about goals, responsibilities, work style, company context, and relevant history. It turns the answers into a small collection of markdown files. Markdown is simply lightweight text formatting that models can read efficiently.
A useful starter folder might contain:
goals.md: What you are trying to accomplish this quarter and year.role.md: Your responsibilities, collaborators, and recurring decisions.company.md: Products, customers, positioning, and important constraints.preferences.md: How you like work presented and where you want pushback.sources.md: A map of important files, apps, and systems of record.CLAUDE.md: The operating instructions for the agent inside that folder.
James uses Obsidian or Cursor to browse those files. Obsidian is a local notes app that makes folders of markdown easier for humans to read and organize.
What goes in CLAUDE.md?
At 19:51, James describes CLAUDE.md as the first file Claude reads in a workspace. It is the map and job description.
A basic version could say:
You are my chief of staff.Before starting work, read goals.md, role.md, company.md, and sources.md.Prioritize tasks that advance the goals in goals.md.Be direct and concise.Flag missing information before making consequential decisions.Never delete files, tasks, events, or records.Ask for approval before sending messages or making financial changes.
The file should point to deeper sources instead of trying to contain your entire life. That keeps the top-level instructions clear and cheaper to load.
Verify the second brain before trusting it
The stream’s first failed demo delivered an important lesson. James had asked a research agent to build background files about Grant and The Neuron. It invented a person and a company.
The right workflow is:
- Let the agent draft the files.
- Read every factual claim.
- Delete hallucinations.
- Add missing context.
- Treat the approved files as source material from then on.
A second brain can concentrate bad information as efficiently as good information. Clean it before building on top.
⚡ Action step: Create four short files: goals.md, role.md, sources.md, and CLAUDE.md. Verify every factual claim before the agent uses them.
Step 3: Connect the tools you already use
Context helps an agent think. Connections let it work.
The standard underneath many of those connections is Model Context Protocol, or MCP. Anthropic introduced MCP as an open way for AI systems to connect with outside tools and data.
James compares an API to a waiter at 25:48. You make a request. The waiter carries it into the kitchen. The result comes back.
MCP gives many AI tools a shared way to work with those waiters. In the interface, you may see “connector,” “plugin,” or “integration.” The underlying concept is the same: the model gains a defined set of actions it can call.
At 30:33, James asks Claude to create a Todoist task for following up with Grant. Claude finds the right project, sets the deadline, and confirms completion. He then changes the priority and adds a note through a second sentence.
The user never writes code. The connector translates the request into the appropriate action.
The copy-paste test
James offers a useful diagnostic at 29:34: repeated copy-paste is a sign that you probably need a connection.
Examples:
- You paste meeting notes into Claude every afternoon. Connect your meeting recorder.
- You paste tasks into Asana or Todoist. Connect the task manager.
- You copy customer details from a CRM into a research prompt. Connect the CRM.
- You paste performance numbers from a dashboard. Connect the underlying data source.
James’s own stack includes Granola for meeting notes, Todoist for tasks, Attio for customer records, Resend for email, GitHub for code, and tools such as Drive, Calendar, Notion, Figma, and Apollo.
The most useful beginner workflow combines two apps:
- Record or transcribe meetings.
- Ask the agent to extract commitments.
- Write those commitments into the task manager.
- Send a short end-of-day summary.
That alone can recover the administrative work hiding between calls.
⚡ Action step: Connect one read-only source. Ask the agent to retrieve a known record, then confirm it found the correct information before granting write access.
Step 4: Design permissions before chasing autonomy
The agent’s power comes from the same thing that makes it risky: it can act across several systems quickly.
James’s first security rule is stronger than “tell the model to be careful.” Remove capabilities it should never use.
At 39:33, he opens Todoist’s connector permissions and blocks the delete tool. He then asks Claude to delete every task in the project. Claude cannot comply because the function is unavailable.
That is better than adding “never delete tasks” to a prompt. Prompts can be misunderstood, overridden, or weakened by conflicting context. Missing permissions cannot be improvised back into existence.
Use three buckets when connecting an app:
- Read: Can the agent search and retrieve information?
- Write: Can it create or update records?
- Delete or execute: Can it remove data, send messages, spend money, run code, or approve transactions?
Start with read access. Add narrow write actions when the workflow requires them. Keep destructive or consequential actions behind human approval.
Anthropic’s own Cowork safety guidance recommends beginning with low-risk tasks, avoiding sensitive or hard-to-reverse actions, and reviewing scheduled outputs. NIST has also identified prompt injection and agent hijacking as active security problems, especially when agents read untrusted emails, websites, and code repositories.
The email problem
James takes a strong position at 1:19:57: connecting a powerful agent directly to your primary inbox creates unnecessary exposure.
Anyone who can email you can place content in front of that agent. A malicious sender could hide instructions inside an ordinary message and attempt to influence another connected tool.
His safer pattern is a private agent inbox:
- Create an inbox whose address stays private.
- Forward or BCC only the messages the agent needs.
- Give the agent access to that curated inbox.
- Require approval before sending replies or changing another system.
Agent Mail provides inboxes designed for agent workflows. The deeper principle is segregation: untrusted public input should not sit next to broad authority.
For highly sensitive work, local models offer another route. Our interview with Intel’s Dr. Olena Zhu shows a local email agent that keeps private data on the device.
⚡ Action step: Open every connector’s permissions. Block deletion, sending, payments, and other irreversible actions unless the workflow truly requires them.
Step 5: Turn repeated work into skills
A good conversation solves the task once. A skill preserves the process.
James defines a skill at 46:39 as a saved instruction package that teaches an agent how to perform a repeatable job.
His Matrix analogy is useful: Neo downloads kung fu, then enters the dojo to test it. The agent gains a capability quickly, but the skill still needs practice and correction.
A skill might cover:
- Turning meeting transcripts into tasks.
- Writing in a specific brand voice.
- Reviewing a product requirements document.
- Running a security checklist.
- Preparing a weekly business review.
- Auditing a website for search problems.
Claude’s official skills documentation describes them as specialized knowledge and workflows. James also shares eight skills he uses personally, and Skills.sh offers a broader public directory.
Do not install random skills blindly
A skill can contain instructions, scripts, and tool calls. Treat it like software.
Grant’s recommendation at 52:53 is practical: give the link to your AI, ask it to inspect the contents, then recreate only the behavior you need.
That approach has three benefits:
- You understand what the skill does.
- You remove unnecessary permissions or code.
- You can adapt the workflow to your own files and tools.
How to create your first skill
James demos the hidden Skill Creator at 53:43. The process is more valuable than the button:
- Define one repeatable task.
- Explain the inputs it should use.
- Specify the exact output format.
- Decide which tools it may call.
- Generate sample outputs.
- Give concrete feedback.
- Save the approved workflow as a skill.
Avoid giant “do everything” skills. A focused skill is easier to test, cheaper to run, and less likely to surprise you.
⚡ Action step: Find a task you have prompted at least twice. Complete it manually with the agent until the output is right, then package that exact process as one focused skill.
Step 6: Test the skill before putting it on a schedule
Most people test a workflow once, see a good result, and automate it. James spends much longer on the quality floor.
At 57:19, he says he may spend 60 to 90 minutes refining a skill he expects to run every day or every week.
Useful feedback sounds like this:
- “This section is too verbose.”
- “That item was trivial; only include strategic commitments.”
- “The report needs direct source links.”
- “Create tasks only when I explicitly made a commitment.”
- “Show the draft invoice, but never send it.”
Then comes the clever part.
James tests skills against historical data or simulated users. At 58:48, he describes spinning up roughly ten subagents, asking each one to run the workflow, and having a master agent inspect the failures.
You can use the same pattern without advanced code:
Simulate ten different users running this skill.Give each user a different role, amount of context, and level of technical ability.Run the workflow for each one.Then analyze the ten outputs and identify:1. Where users became confused.2. Which instructions produced inconsistent results.3. Which permissions or assumptions created risk.4. How the skill should change before I use it in production.
Simulation cannot prove a workflow is safe. It can expose obvious weak spots before the agent reaches real customers, files, or money.
⚡ Action step: Run the skill against at least three past examples and one deliberately messy example. Fix every failure you can reproduce.
Step 7: Put a reliable skill on a schedule
A skill becomes proactive when something triggers it.
The simplest trigger is time. Claude Cowork’s scheduled tasks can run reports, research, and connected workflows on a recurring cadence.
James separates scheduled work into two categories at 1:01:45:
Local scheduled tasks
These run on your computer.
Advantages:
- They can reach local files and folders.
- They can work with software installed on the machine.
- Your second brain can remain local.
Tradeoff:
- The computer must remain on.
- The relevant application must remain available.
Cloud scheduled tasks
These run remotely while the laptop is asleep or turned off.
Advantages:
- They are always available.
- They can use cloud connectors and account-hosted files.
- They work well for daily briefings and recurring reports.
Tradeoff:
- They cannot read a file that exists only on your laptop.
- The required apps and connectors must be authenticated.
That storage decision matters. A second brain held only in Obsidian on your computer works well for local tasks. A cloud agent may need the same context in Notion, Google Drive, or another connected source.
The chief-of-staff workflow
James attempts to build a cloud chief of staff at 1:04:41. The intended workflow is:
- Read the day’s meeting transcripts.
- Read the user’s goals and role context.
- Extract commitments through a saved skill.
- Create the relevant Todoist tasks.
- Email a concise summary.
The live run fails because some connectors are not authenticated and the chosen skill points toward Granola instead of the demo transcripts in Notion.
That failure is useful. Scheduled tasks depend on the boring details:
- Is every connector authenticated?
- Can the agent reach the exact file location?
- Does the skill reference the correct source?
- Does the selected model have enough capacity?
- Does the workflow stop safely when information is missing?
The production version works for James. His daily agent summarizes meetings and creates five commitments. Other agents review his sales pipeline and propose website fixes through GitHub.
⚡ Action step: Schedule the workflow at a conservative cadence and review every run. Weekly is safer than hourly while the system is still learning the job.
Step 8: Move from reporting to proactive work
A daily summary feels impressive for a week. Then it becomes another message you ignore.
James’s four-level framework at 1:12:05 separates useful automation from inbox furniture.
Level 1: Report
The agent tells you what happened.
Example: “You had three sales calls. Two prospects requested pricing. One deal has been idle for 12 days.”
Level 2: Recommend
The agent uses goals and context to suggest where you should focus.
Example: “Follow up with these two prospects first because they match your target customer and represent the highest potential revenue.”
Level 3: Draft the work
The agent prepares the next action.
Example: It drafts both follow-up emails, updates the deal stages, and prepares the invoices for approval.
Level 4: Improve itself
The agent reviews its past outputs and adjusts the workflow.
Example: It notices that you repeatedly ignore one section of the briefing, removes that section, and updates its instructions.
Level four needs guardrails. An agent that edits its own instructions can also weaken a safety rule or optimize toward the wrong signal. Keep protected instructions, change logs, tests, and human approval around self-modification.
The direction is still right: useful agents should learn from feedback instead of repeating the same mediocre report forever.
⚡ Action step: Upgrade one existing report so the agent also recommends a next step or drafts the work for approval.
Step 9: Choose the right model for each part of the job
Agents create a new cost problem. The best model can perform every step, but many steps do not require the best model.
James’s rule at 1:27:30 is simple:
- Use a “big brain” model for planning, designing the skill, and resolving ambiguity.
- Use a cheaper worker model for structured execution.
A weekly report skill may need a strong model during creation. Once the instructions are precise, a faster model can collect the data and format the output.
Grant applies the same logic to larger projects. A high-end model writes the technical plan. Lower-cost subagents implement pieces of it. Another model reviews the combined result.
OpenRouter gives users one interface for trying models from many providers. Artificial Analysis compares model quality, speed, and pricing.
James warns against routing every Claude Code request through an API without understanding the billing. You can accidentally replace a predictable subscription with expensive usage charges.
Start with the subscription you already have. Add model routing after you can measure which steps are expensive and which steps deserve more intelligence.
⚡ Action step: Use the strongest model to design the workflow, then test whether a cheaper model can run it without lowering the quality floor.
Step 10: Scale the knowledge system, not the context window
A large context window is not a filing system.
When a business has hundreds of thousands of emails and documents, the agent should find the small relevant subset instead of loading everything.
At 44:14, James recommends:
- Index pages that explain where information lives.
- Short summaries at the top of long documents.
- Links between related files.
- Clear names and folder structures.
- A source-of-truth map in the agent’s instructions.
Andrej Karpathy’s LLM Wiki pattern offers one approach to making a document collection easier for models to navigate.
At enterprise scale, a dedicated search layer such as Glean may become necessary. The core idea remains the same: retrieve the right evidence first, then reason over it.
⚡ Action step: Create one index page that tells the agent where your important sources live, what each source contains, and which source wins when two documents disagree.
What this guide cannot automate away
The agent story often gets marketed as “say what you want, then go drink coffee.” The stream shows the less glamorous reality.
The context files require review. Connectors require authentication. Permissions require judgment. Skills require testing. Schedules require monitoring. Models behave differently. Cloud and local storage create tradeoffs.
The live demos fail twice: the context demo stumbles at 22:39, and the scheduled chief-of-staff run breaks at 1:08:11.
That does not disprove the value. It identifies the actual skill: agent building is operational design. You are deciding which information enters the system, which actions it can take, how success is measured, and where a human remains responsible.
NIST’s 2026 work on agent security found broad agreement that these systems introduce novel threats and that ordinary cybersecurity controls need adaptation. Anthropic tells users to begin with low-risk scheduled tasks and avoid sensitive, consequential actions.
The most credible version of the agent future includes more controls, testing, and audit logs, not fewer.
Advanced path: Deploy the agent as a cloud service
The final demo points toward where agent building may go next.
Vercel Eve treats an agent as a directory. Markdown files define its instructions and skills. TypeScript files define its tools. The framework handles durable execution, schedules, connections, and channels.
At 1:53:34, James shows a health workspace connected to Garmin, TrainingPeaks, blood work, and other files. He uses Eve to deploy it as a Telegram agent in about 30 minutes.
The result sends:
- A morning health briefing.
- Midday nudges based on nutrition and training data.
- Answers through Telegram.
- Cloud-hosted execution without a dedicated computer.
The useful part is timing. A dashboard waits for the user to remember it. A message can intervene while the decision is still being made.
The remaining gap is product design. Beginners still need a clean setup flow for model credentials, hosting, channels, permissions, and data sources. The tool that hides those mechanics must also make the risks visible.
That tension will decide whether consumer agents become ordinary software or a series of powerful experiments maintained by enthusiasts.
Your action plan for the next seven days
So I promised at 1:58:10 to turn the two-hour session into an easier sequence. Here it is. Start with one narrow workflow that removes real work. Five agents can wait.
Today: Choose the task
Pick something that happens at least weekly and has a clear output.
Good first candidates:
- Extract commitments from meetings.
- Prepare a daily schedule.
- Draft a weekly status report.
- Research new leads.
- Review a content calendar.
- Summarize a small private inbox.
Avoid payments, deletions, public posting, legal decisions, or unrestricted email at the beginning.
Tomorrow: Build the context folder
Create goals.md, role.md, sources.md, and CLAUDE.md. Keep each file short. Verify every fact.
Day 3: Connect one read-only source
Connect the app that contains the necessary input. Test simple searches before adding write access.
Day 4: Run the workflow manually
Complete the job in conversation. Correct the result until the output is useful in real work.
Day 5: Turn it into a skill
Package the approved instructions, sources, tools, and output format.
Day 6: Test it
Run historical examples. Simulate edge cases. Remove assumptions. Confirm the workflow stops when information is missing.
Day 7: Schedule it
Choose a conservative cadence. Review every run. Add permissions only when the agent has earned them.
The system compounds in that order. Context is onboarding. Connections provide access. Skills provide training. Schedules provide initiative.
Or, if you have some time on your hands, do all of those in one afternoon!
Full key moments from the livestream
The TL;DR above is the fastest route. This is the complete 143-point index from our video-insights pass, preserved here so you can jump to any claim, demo, warning, or audience question without scrubbing through two hours of video.
Foundations: what an agent is and why context comes first
- (02:02) Grant introduces James McAulay as a former ElevenLabs operator who helped the company grow from roughly $110M to $300M in annual recurring revenue, then built an AI-native business that reached about $80K in monthly revenue by month three and a $200K-plus month by month five.
- (02:27) James says a major reason the new business grew faster than expected was that he structured the company around agents from day one, with himself and Claude Code doing the work rather than a traditional employee base.
- (03:50) James frames the session around four fundamentals, then promises to combine them into a proactive chief-of-staff agent that can work without waiting for a fresh prompt.
- (04:26) Cloud-scheduled tasks remove the old requirement to leave a computer running around the clock, which had pushed some users to buy dedicated Mac minis.
- (05:19) James defines an agent as three things working together: an LLM “brain,” context about the user or business, and tools that let it take action.
- (05:52) His core shift is from talking about work with a chatbot to delegating work to a system that returns with the report finished, the email drafted, or the task completed.
- (06:38) James avoids relying on automatic memory because it remembers fragments. He prefers explicit documents and system instructions that the agent can read repeatedly.
- (07:11) Useful agent context can come from static text files, behavioral instructions, and live sources such as Google Drive or Notion.
- (08:10) The “CFO at coffee” analogy explains why context matters: a brilliant CFO with no business data can only offer generic advice or spend the whole meeting asking questions.
- (09:08) Give that same CFO read-only access to Stripe, Xero, or QuickBooks first, and the conversation can begin with a specific diagnosis instead of textbook advice.
- (09:56) James compares an out-of-the-box chatbot to a coach with bad memory: every useful conversation requires bringing the flight logs again, and the coach still forgets most of the last session.
- (10:43) The target experience is a co-pilot that has been in the plane with you, sees the live instruments and history, can update the flight plan, and handles checks while the human still flies.
- (11:30) Custom agents once required teams of engineers and months of development. James argues that businesses can now build them inside frontier-model ecosystems without a technical background.
- (12:14) For beginners, James recommends Claude Cowork as a lower-friction and slightly lower-risk bridge between ordinary chat and Claude Code.
- (13:31) The session’s one-sentence definition: an agent has an LLM brain, context, and the ability to do things.
- (14:04) James moved from a $20 Claude plan to roughly $90 and then $200, but argues the Max 5x plan can return far more than its cost when agents perform real business work.
- (14:41) His “Set Up My Second Brain” skill conducts a two- to three-hour interview about goals, responsibilities, work style, LinkedIn history, and company context, then turns the answers into a compact folder of files.
- (15:10) James treats this small context folder as the beginning of a second brain: a portable packet that helps any compatible agent understand who you are, what you want, and how you work.
- (15:45) After running the exercise with about 400 people, James says the jump from partial familiarity to rich context produces dramatically better outputs.
- (16:31) He uses Obsidian or Cursor to browse and manage these markdown context files.
- (17:13) Markdown works well for agent context because it is lightweight, compact, and easy for models to parse.
- (18:11) James recommends asking Claude Research to investigate you and your company, producing a deeper background report from hundreds of sources.
- (18:53) The demo also reveals a warning: research agents can hallucinate people and companies, so generated background files still need human verification.
- (19:22) A role-profile file can capture responsibilities, collaborators, expertise, and prior experience, allowing the agent to connect a current task with work you have already done elsewhere.
- (19:51) Claude.md is the first file Claude reads in a folder. It can assign a role, point to goals and supporting files, define behavior, and tell the agent how to follow up.
- (20:30) James describes Claude.md as the map of the workspace: the file that tells any arriving agent who it is, what the folder contains, and what outcome it should pursue.
- (21:19) When Cowork opens the chief-of-staff folder, it can read the included context before answering routine questions such as “help me plan my day.”
- (22:09) The practical promise resembles J.A.R.V.I.S.: instead of explaining who “James” is every time, the agent should infer the relevant person, project, and situation from stored context.
- (22:39) The failed live demo exposes a model-selection lesson: Sonnet can optimize too aggressively for speed and brevity when the task requires deeper context synthesis.
- (23:23) Grant recommends Opus or Fable over Sonnet when the task demands the fullest possible use of a large context.
- (24:22) Account-level memory and folder-level instructions can conflict. Users need to remember that the model may combine both unless one is disabled or clarified.
- (24:49) For a long context interview, James recommends dictation software such as Wispr Flow because agents are good at turning unstructured spoken thinking into structured files.
Connections, permissions, and organizing business knowledge
- (25:48) James explains an API as the waiter between two software systems: it carries a request into the kitchen and brings the result back.
- (26:47) Before MCP, each AI system built its own custom way to talk to every API, forcing agents to improvise code and vendors to maintain duplicate integrations.
- (27:28) Model Context Protocol standardizes that connection so a vendor can expose one interface that many agents can use.
- (28:12) In plain English, MCP translates a user request into code, sends it to the connected service, then converts the response back into language.
- (28:56) Grant grounds MCP for beginners: the connector or plugin buttons inside ChatGPT and Claude are the visible layer; MCP is often the plumbing underneath.
- (29:34) A recurring sign that you need an integration is manual copy-paste. If you keep pasting information from one app into an agent, connect the source directly.
- (30:33) The Todoist demo shows natural-language task creation: Claude finds the right project, creates a follow-up task, sets a deadline, and returns a confirmation.
- (31:56) The task can be edited conversationally too, including priority, labels, and notes about the specific reason for following up.
- (32:43) James increasingly avoids opening task-management apps. He asks Claude what is due, what needs attention, and which tasks the agent itself can take over.
- (33:30) He describes two maturity levels: first the agent can see human work; next the human can delegate items directly from the shared task board.
- (34:40) A high-value starter stack is task management plus meeting transcription. At day’s end, the agent can scan calls, identify commitments, and create the relevant tasks.
- (35:16) Grant summarizes the beginner setup: start with a paid AI plan, use an action-capable interface such as Cowork, connect the apps you already use, then delegate in plain language.
- (35:53) James connects Apollo for contact enrichment, Figma for diagrams and mockups, GitHub for his site, Calendar, Drive, Granola, Notion, Resend, Todoist, Attio, and other systems.
- (37:15) A cross-app sales workflow can review meeting transcripts, update CRM stages, create follow-up tasks, and prepare invoices for human approval.
- (38:15) The result is a virtual sales assistant that attends every meeting, maintains the CRM, tracks commitments, prepares invoices, and can flag overdue payments.
- (38:44) James now treats agent compatibility as a purchasing requirement: if a SaaS product cannot connect to Claude, he will look for an alternative.
- (39:33) His first security rule is permission design. Prevent destructive actions by removing the capability itself, rather than merely instructing the model not to use it.
- (40:06) Connector permissions should be reviewed tool by tool. Read access, write access, and delete access carry different risk and should not be bundled casually.
- (41:21) Blocking the Todoist delete tool proves the distinction: even when asked to delete every task, Claude cannot comply because the capability no longer exists.
- (41:35) Natural-language ambiguity creates risk. “Tidy up my to-do list” might be interpreted as deleting items, so destructive permissions should remain unavailable by default.
- (42:23) Personal-account users should disable model-improvement data sharing when handling sensitive work. James says Team and Enterprise plans disable training on customer data by default.
- (42:52) Data-residency rules may still block some cloud deployments, especially for organizations that cannot move data across regions. Enterprise architecture may require AWS or other controlled hosting.
- (43:13) Grant points to Dr. Olena Zhu’s local email agent as a privacy-oriented alternative for companies that cannot send confidential email through a cloud model.
- (43:35) James predicts smaller local models will handle narrow tasks such as email summarization because those jobs do not require the most capable frontier model.
- (44:14) For very large knowledge bases, James recommends indexes that explain where files live, summaries at the top of documents, and links between related files.
- (44:52) Karpathy’s “LLM Wiki” provides a pattern for organizing knowledge so agents can navigate it with fewer tokens and less search.
- (45:04) At larger scale, companies may need a dedicated enterprise search layer such as Glean to retrieve relevant information from millions of records.
- (45:54) Grant’s version of the same architecture is a source-of-truth map in project instructions or skills, telling the agent which files exist and when to reference them.
Skills: creating, refining, and testing repeatable work
- (46:39) James defines a skill as a saved instruction package that teaches an agent how to perform a repeatable job.
- (47:09) He uses Neo learning kung fu in The Matrix as the metaphor: install a skill quickly, then test and refine it in a controlled environment.
- (48:49) Skills can trigger automatically from the user’s request, create more consistent outputs than freestyle prompting, improve over time, and be shared with teammates or customers.
- (49:47) James’s examples include meeting-to-task extraction, writing humanization, brand voice, security checks, and product-requirement documents.
- (51:29) Skills.sh is presented as a large directory of installable skills, including marketing psychology and SEO-audit workflows.
- (52:53) Grant’s safety advice is to give a third-party skill link to Claude, ask it to inspect the contents, and recreate the useful behavior rather than blindly installing something suspicious.
- (53:43) The hidden Skill Creator interviews the user, asks about output format and downstream actions, generates examples, requests feedback, and packages the result into a reusable skill.
- (55:39) A well-built skill should specify whether to save a report, create tasks, use bullets or paragraphs, and how much detail the output should contain.
- (56:14) The purpose of packaging a workflow as a skill is predictability: the same job should produce a similar structure and quality each time.
- (57:19) James may spend 60 to 90 minutes refining a skill he expects to use every day or every week.
- (57:47) Quality improves when feedback is concrete: identify what was too verbose, what was trivial, what should count as strategic, and what the model should ignore.
- (58:17) For a higher quality floor, James simulates the skill on historical days or weeks rather than waiting to discover flaws in live use.
- (58:48) His advanced test is to spin up roughly 10 subagents, have each simulate a different user or time period, then ask a master agent to analyze where the workflow failed.
- (59:54) Grant contrasts this with his own more iterative method: use the skill in real work, point out mistakes, and add “never make this mistake again” instructions over time.
- (01:00:37) Cowork can launch subagents when explicitly instructed, giving nontechnical users a way to test a workflow across several synthetic personas.
Scheduled tasks and the four levels of proactive agents
- (01:01:45) James maps scheduled work into two dimensions: cloud versus local, and Cowork versus Code.
- (01:02:15) Local scheduled tasks can access files on the computer but require the machine to remain on and Claude to stay open.
- (01:03:14) Cloud-scheduled tasks run on Anthropic’s infrastructure while the computer is off, but they cannot read local files.
- (01:03:42) That limitation makes cloud-hosted context stores such as Notion more useful for always-on agents than a second brain that exists only on a laptop.
- (01:04:41) A cloud chief-of-staff task can read Notion goals and meeting transcripts, use a skill to extract commitments, write tasks into Todoist, and email a summary.
- (01:06:03) The scheduling interface can set the model, recurrence, permissions, and whether the task runs while the computer is off.
- (01:08:11) The failed demo is itself a setup lesson: scheduled agents break when connectors are unauthenticated or the wrong skill points to the wrong source.
- (01:08:53) James says his production version runs daily, summarizes meetings, identifies five commitments, and inserts them into Todoist.
- (01:09:22) Other recurring agents include a weekly SEO agent that proposes GitHub pull requests and a morning sales agent that briefs him on pipeline and upcoming calls.
- (01:10:32) A year earlier, James would not have known how to build an agent with rich context, connected tools, and a 7 PM schedule. He now sees it as an approachable configuration problem.
- (01:12:05) James presents four levels of proactive agents, beginning with a system that simply reports what happened.
- (01:12:34) Level two adds recommendations, such as which opportunities deserve attention based on revenue potential and ideal-customer fit.
- (01:13:03) Level three drafts the work itself: follow-up emails, invoices, or other deliverables that a human only needs to review and approve.
- (01:13:47) Level four evaluates its own past outputs, notices repeated low-value behavior, and updates its instructions so performance compounds.
- (01:14:30) James warns that briefing-only agents quickly become background noise. Proactive agents stay useful by completing work and improving their own process.
- (01:16:47) A Claude.md instruction can ask the agent to notice repetitive work and suggest creating a new skill whenever it sees a recurring pattern.
- (01:17:10) James’s weekly Workspace Review skill audits context, finds missing files, proposes new skills or connectors, and updates Claude.md based on repeated feedback.
Email safety, remote access, model choice, and local AI
- (01:19:57) James strongly advises against connecting an all-powerful agent directly to a primary Gmail inbox because any sender can potentially place instructions in front of the agent.
- (01:20:27) His prompt-injection example hides malicious instructions in white text inside an ordinary-looking email, hoping the connected agent will obey them and use another tool such as accounting software.
- (01:21:23) The safer pattern is a private agent inbox whose address nobody else knows. The human forwards or BCCs only the messages the agent should see.
- (01:22:22) Agent Mail gives that private inbox an MCP connection, letting a morning stand-up scan a small, curated queue instead of a chaotic primary mailbox.
- (01:24:28) For remote Claude Code use, Grant recommends starting the session locally, enabling Remote Control, and then continuing from a phone while the home computer remains on.
- (01:25:03) James previously used a Telegram bridge to pipe mobile messages into a Claude Code terminal, but found the setup secure yet fiddly.
- (01:26:29) Claude Code’s built-in Remote Control now offers the simpler route: enable it in the terminal or app, then open the same live session on mobile.
- (01:27:30) James uses the strongest model, such as Fable or Opus, to create skills, then runs simple skills with Sonnet or even Haiku to save tokens.
- (01:28:17) His general model rule is “big brain” for planning and quality-sensitive design, cheaper worker models for deterministic execution.
- (01:28:40) OpenRouter lets users access models from Anthropic, OpenAI, Meta, DeepSeek, Qwen, Kimi, GLM, and others through one interface.
- (01:29:09) Routing all Claude Code traffic through OpenRouter can accidentally replace the Max subscription with expensive API billing.
- (01:29:47) James built custom commands that keep Claude on the subscription while sending only Kimi or GLM calls through paid APIs.
- (01:30:23) He says GLM was fast and found an issue that Fable missed, reinforcing the value of testing cheaper models instead of assuming the flagship always wins.
- (01:31:15) Grant’s model-effort framework separates intelligence from time: raise the model and effort when a task requires broad context, difficult reasoning, or sustained work.
- (01:31:43) Both speakers warn that extreme effort modes can trigger huge swarms of subagents and burn through usage limits without proportional benefit.
- (01:32:39) Grant uses Fable as a planner that writes a detailed technical specification, then delegates execution to lower-cost subagents.
- (01:33:07) James adds a multi-agent refinement pass: several agents critique a product-requirements document from different perspectives before workers implement it.
- (01:34:12) Cloud tasks can read Google-hosted context while the computer is off, but James considers Notion more dependable for workflows that need to create or update cloud documents.
- (01:35:18) James criticizes Gemini’s closed ecosystem: trainees often abandon it because they cannot connect the broader set of tools their workflows require.
- (01:36:15) A correction from the chat notes that Claude can write markdown files to Google Drive, showing why live technical claims should be checked against actual connector behavior.
- (01:37:20) James does not maintain many separate “agents.” He thinks of Claude Code as one agent equipped with many skills.
- (01:37:49) When a new model arrives, he reevaluates the skills and instructions because stronger models may need less hand-holding than older ones.
- (01:38:45) Usage can be monitored inside Claude, while the Claude Code HUD plugin shows the current model, context consumption, and reset window directly in the terminal.
- (01:40:46) James concludes that running giant local models at roughly two tokens per second is often less practical than paying pennies for cloud inference.
- (01:41:17) After buying a 64 GB Mac mini for local inference, he shifted toward OpenRouter because downloading, configuring, and running large models was more trouble than the savings justified.
- (01:41:49) Grant remains bullish on local AI and cites Dr. Olena Zhu’s forecast that Fable-class capability could reach laptops within roughly two years if current trends hold.
- (01:42:17) James agrees that local or privately hosted models are the cleanest route for organizations that need strong guarantees their data will not leave controlled infrastructure.
- (01:42:42) Artificial Analysis can help users compare hosting providers, model quality, price, and private deployment options.
- (01:43:03) OpenRouter Fusion is described as a possible model router that selects or combines multiple models to produce a panel of answers.
Triggers, GitHub, portability, and cloud deployment
- (01:43:55) For advertising workflows, James mentions Meta’s CLI and Windsor.ai as ways to connect Google, Meta, Shopify, and other data through a single MCP layer.
- (01:44:21) One AI-native growth agency pipes ad data into BigQuery, exposes Python scripts to agents, and runs scheduled checks that flag creative fatigue in Slack.
- (01:45:21) Trigger-based agents can run from a webhook rather than a clock: an API call to a Claude Code routine can launch a connected workflow.
- (01:45:51) Cloud routines that execute against code require a GitHub repository, while local routines can run directly on the computer.
- (01:47:03) James explains GitHub for beginners as a cloud home for code that supports collaboration, private repositories, open source, and disposable cloud execution environments.
- (01:47:52) Despite studying computer science, James says he has not run a Git command manually in a year; he asks Claude to create branches, run tests, and submit pull requests.
- (01:48:26) Parallel coding agents collide when they edit the same files. Git worktrees give each agent an isolated copy so one worker does not overwrite another’s changes.
- (01:49:12) Grant argues the next step is moving parallel agent work from local worktrees into cloud environments so duplicated repositories do not fill the laptop’s storage.
- (01:51:14) The concepts transfer across platforms: markdown context, MCP connectors, and skills are portable even when the exact scheduling interface differs.
- (01:51:52) ChatGPT workspace agents may be easier to configure than Claude scheduled tasks, but Grant hit the monthly workspace-credit ceiling after roughly one week of heavy automation.
- (01:53:01) That scarcity strengthens the case for hybrid architectures where scheduled work can use local or inexpensive models through independent harnesses such as Hermes or OpenClaw.
- (01:53:34) James introduces Vercel Eve as a new framework that can turn a folder into a cloud-hosted agent reachable through Telegram, Slack, WhatsApp, or a website.
- (01:54:02) He used Eve inside a health-and-fitness workspace connected to Garmin, TrainingPeaks, blood work, and other files.
- (01:54:59) Within about 30 minutes, he had a Telegram health bot that sends a morning briefing and midday behavioral nudges based on nutrition and training data.
- (01:55:11) The agent works because timely messages change behavior better than a dashboard a person has to remember to open.
- (01:55:57) The health bot runs in Vercel’s cloud, can switch models, and does not require a VPS, a dedicated Mac mini, or a computer left on.
- (01:56:07) Grant identifies the remaining adoption gap: beginners still need a product that abstracts endpoints, gateways, app tokens, hosting, and channels into one setup flow.
- (01:56:35) James thinks Eve is close to that experience, especially for Telegram, because it reduced deployment to roughly half an hour.
- (01:57:33) The first 24 hours of health-bot testing cost James about $2 on Sonnet 5 for roughly 40 to 50 messages.
- (01:57:45) He expects a cheaper model such as DeepSeek or GLM to reduce ongoing costs to pennies, and considers even $10 per month worthwhile if the agent measurably changes his behavior.
- (01:58:10) Grant promises to turn the transcript into a structured article that answers the unanswered chat questions and puts the two-hour lesson into an easier sequence.
- (01:58:32) James points viewers to Agent Accelerator, LinkedIn, and his YouTube channel, where he planned a tutorial on running Kimi and GLM inside Claude Code.
Now let’s turn those 143 moments into one workflow you can follow from beginning to end.
BONUS Q&A: Questions from the live chat
“What is the relationship between a chat, a project, and an agent?”
At 13:31, James reduces the agent to its three working parts. A chat is one conversation. A project is a workspace containing shared files, instructions, memory, and related conversations. An agent is the system using that context and tools to pursue a goal or complete work. We also go over the difference between projects and skills and chats in this video, so you might want to watch that for a refresher.
“Do I feed the folder to Claude every time?”
At 21:19, James opens the chief-of-staff folder directly in Cowork. Claude reads the main instruction file and follows references to the other files as needed. You do not manually upload every file on every turn.
“Can I do this in ChatGPT, Copilot, or Gemini?”
The framework transfers across platforms, as Grant and James discuss at 1:51:14: persistent context, tool connections, reusable workflows, permissions, and triggers. The interfaces differ. James teaches primarily in Claude because he currently prefers its agent tooling, but the underlying skill is platform-agnostic.
Our earlier AI for Total Beginners companion guide compares projects, skills, and automations across the major platforms. The original video starts its framework at 4:03.
“Should I start with Notion or Obsidian for a second brain?”
James uses Obsidian or Cursor for local markdown at 16:31, then explains why cloud tasks favor Notion at 1:03:42. Use Obsidian for portability and local privacy. Use Notion when cloud tasks need to reach and update the same knowledge while your computer is off.
“Can the agent search 100,000 emails or files without running out of context?”
At 44:14, James recommends retrieval and navigation rather than loading the whole collection. Add indexes, summaries, links, and a map of sources. Larger organizations may need enterprise search infrastructure.
“How do I know which model is right for a task?”
James gives the model-selection rule at 1:27:30: use the strongest model for ambiguous planning, important reasoning, or skill design. Use cheaper models for repeatable execution after the workflow is well specified.
“How often should I reevaluate the agent?”
James says at 1:37:49 that new models can justify revisiting skills and instructions. Review early runs closely, then schedule a weekly or monthly workspace review for repeated failures, missing context, and unused outputs.
“How can I see token usage?”
At 1:38:45, James shows that Claude reports usage inside the app. He also uses a Claude Code HUD plugin that displays the current model, context usage, and reset window inside the terminal.
“Can I give the agent access to Gmail?”
You can, but James warns against broad primary-inbox access at 1:19:57. A private agent inbox limits untrusted input. Keep permissions narrow and require approval before sending or changing consequential systems.
“What are the best meeting note tools?”
James recommends meeting transcription plus task management at 34:40 and uses Granola in his own stack. The important feature is a reliable transcript the agent can access.
“Can agents run from a trigger instead of a schedule?”
Yes. At 1:45:21, James shows that a webhook or API event can launch a workflow when something happens. Triggered routines require more setup than a simple schedule.
“How do I use Claude Code while traveling?”
At 1:24:28, Grant recommends starting the session on the home computer, enabling Remote Control, and continuing from the phone while the machine stays on.
James previously used a Telegram bridge, but the built-in remote option is simpler.
“Can I use local models instead of paying for cloud inference?”
Yes, especially for privacy-sensitive or narrow tasks. At 1:40:46, James says very large local models running slowly were less practical for him than cheap cloud inference.
“How transferable are agent skills between platforms?”
At 1:51:14, Grant and James agree that markdown instructions, workflow logic, examples, and safety rules transfer well. Tool names, authentication, memory, scheduling, and file access require platform-specific changes.
“Should I create many agents or one agent with many skills?”
James explains at 1:37:20 that he treats Claude Code as the main agent and equips it with many focused skills. Separate agents become useful when they need different permissions, identities, or isolated contexts.
“What is the best first business agent?”
Start with the workflow James recommends at 34:40: read meeting transcripts, organize commitments, and draft or create tasks for review. The inputs and expected result are easy to inspect.
All resources and links
Watch and learn
- Full AI agents livestream with James McAulay
- The Neuron YouTube channel
- AI Agents Are About to Move Off the Cloud with Intel’s Dr. Olena Zhu
- The 5-Step Framework to Learn AI in 2026
James McAulay
James shares his course, social profiles, and YouTube tutorials at 1:58:32.
- Agentic Growth Accelerator
- The eight Claude skills James uses
- James McAulay on YouTube
- James McAulay on LinkedIn
Second brains and meeting context
Skills, connectors, and models
Proactive agent infrastructure
Keep learning
- AI for Total Beginners companion guide
- How to Actually Use AI in 2026: The Complete Guide
- The Neuron Academy
- Subscribe to The Neuron
The next breakthrough in beginner agents will not come from a model suddenly becoming perfect. It will come from a product that makes context, permissions, testing, and deployment feel obvious without hiding the consequences.
When that product arrives, the difficult question changes. Building an agent becomes easy. Deciding what authority it deserves remains a human job.