There is nothing more annoying than being in the middle of an important task with an AI assistant and suddenly seeing a message telling you that you have reached your usage limit. You could be debugging code. You could be writing an article. You could be analyzing a document. You could be researching something for work. And just when the AI finally understands what you are trying to do, you are told to come back later.
The obvious reaction is to think: “I need a better plan.”
But sometimes the problem isn’t the plan you are paying for. It is the way you are using it.
AI usage limits aren’t simply about the number of messages you send. The amount of context in those messages, the length of the conversation, the model you use, the effort level, attached files, and tools such as web research can all affect consumption. Anthropic explicitly lists message length, attachment size, conversation length, tools, model choice, effort level, and multi-step tasks among the factors affecting Claude usage.
/That means there are some surprisingly simple changes you can make to your workflow. You don’t need to send fewer prompts. You need to make each prompt and conversation more efficient.
Here is the workflow I recommend.
1. Stop Treating Every AI Chat Like One Endless Conversation
This is probably the easiest habit to change. When we start using an AI assistant for something, it is tempting to keep the same chat open forever.
You ask a question. Then a follow-up. Then another. Then you switch topics. Then you remember something from yesterday. Then you start another task. Eventually, you have a massive conversation containing dozens of unrelated questions. It feels convenient because everything is in one place. But it isn’t necessarily efficient.
The longer a conversation becomes, the more context the AI has to work with. Anthropic specifically recommends starting a new conversation when you are approaching usage limits in a long chat, and its documentation distinguishes usage limits from context-window limits.
My simple rule: One meaningful task = one chat.
For example:
- Chat 1: Fix the authentication bug
- Chat 2: Review this Python script
- Chat 3: Write documentation for the API
- Chat 4: Research a new laptop
- Chat 5: Draft a blog post
Don’t turn Chat 1 into: Fix authentication → explain Docker → help me write an email → what’s the best MacBook → now review this code.
Starting a new chat isn’t losing context. It is removing context you no longer need.

2. But Don’t Start a New Chat When Context Actually Matters
There is an important exception. If you’re working on the same project, repeatedly starting from zero can be counterproductive. Suppose you’re building an application. The AI already knows:
- The architecture
- The database structure
- The naming conventions
- The API design
- The existing bugs
- Your preferred coding style
Starting a completely new chat means you now have to explain all of that again. So the trick isn’t: “Always start a new chat.” It is: “Start a new chat when the old context isn’t useful.”
For larger ongoing projects, use project-level knowledge, memory, or a concise project document where supported. The goal is to keep persistent context separate from temporary conversation history.
3. Give the AI Enough Context — But Not Everything
This is probably the biggest mistake people make. When an AI asks for context, many people respond by throwing the entire universe at it. Entire repositories. Entire PDFs. Entire log files.
The AI might be capable of handling it, but that doesn’t mean it is an efficient way to use your usage allowance. A better approach is to ask: What does the AI actually need to answer this question?
If you’re debugging an error, perhaps it needs the error message, the relevant function, and the surrounding configuration. It probably doesn’t need your entire repository. If you’re asking about a contract clause, give it the clause and the surrounding context instead of the entire 40-page contract.
Think in terms of minimum useful context. Not: “Here is everything.” But: “Here is everything you need.”
4. Don’t Upload a Huge File If You Only Need One Small Part
Files are particularly easy to overlook. You might have a 30 MB PDF and think: “The AI can read PDFs, so I’ll just upload it.”
Large attachments can consume significantly more context than a short piece of extracted text. Imagine you have a 100-page technical manual. You need to know: “What does the reset procedure on page 47 say?” Uploading the entire manual is unnecessary if you already know exactly what you’re looking for. Extract the relevant section. Give the AI that. Done.
One useful rule: If you can describe exactly where the answer is, don’t make the AI search through everything else.
5. Screenshots Can Be Useful — But Crop Them
Screenshots are incredibly useful when working with AI, especially for error messages, UI problems, dashboards, or application settings. But don’t screenshot your entire 4K monitor when the important information occupies a tiny corner.
Crop it. If the error is a tiny box in the center of your screen, don’t send the browser, the taskbar, your Slack messages, and the terminal background. Send the error. A tightly cropped image is easier for you to inspect, easier for the model to understand, and potentially much lighter on your context limits.
6. Write Better Prompts Instead of Sending Five Follow-Ups
This is where a lot of usage gets wasted.
You ask: “Write a blog post about AI.” AI responds. You say: “Make it more interesting.” AI responds. “Make it longer.” AI responds. “Add SEO.” AI responds.
You’ve now used six messages to arrive at something you could have described in one reasonably detailed prompt. A better prompt gives the model the destination before it starts:
“Write an 8-minute conversational technology article about managing AI usage limits. Explain context windows, new chats, file sizes, focused prompts, and model selection. Include examples and a short conclusion.”
The goal isn’t to write enormous prompts. It is to avoid incremental prompting caused by an incomplete first prompt.
7. Batch Related Requests
If you have five related things to ask, don’t necessarily send five messages. Instead of asking it to explain the code, then finding bugs, then suggesting improvements, then writing tests—batch them together.
“Review this code. Identify bugs, suggest improvements, explain the important sections, then provide a test plan. Prioritize correctness over style.”
Now the AI has the entire task. Anthropic’s usage guidance explicitly recommends batching similar requests into one message.
8. Use the Right Model for the Job
Not every task deserves your most capable model. Do you really need the highest reasoning model to rewrite an email, fix grammar, or convert JSON to YAML? Probably not.
Save the heavier, token-hungry models for tasks that genuinely benefit from deeper reasoning: complex debugging, architecture decisions, difficult research, or multi-step agentic tasks.
Simple principle: Don’t use a Formula 1 car to go to the grocery store.
9. Don’t Ask the AI to Rewrite the Entire Thing When One Section Is Wrong
Imagine the AI writes a 2,000-word article, but you don’t like section 4.
Instead of saying: “Rewrite section 4. Keep everything else unchanged,” you say: “Rewrite the article.” Now the model generates the entire 2,000 words again.
If only one component is wrong, fix that component. Identify the specific section that needs changing rather than asking the model to regenerate the entire output.
10. Tell the AI Exactly Where to Look
This becomes particularly important with coding agents and connected tools. If you’re working inside a large project and you know the relevant file, tell the AI.
Instead of: “Find where authentication is implemented.” Try: “The issue is probably in auth/session.ts. Check that file and the middleware it calls.”
Pointing an agent directly at the relevant folder or file rather than making it search through everything saves massive amounts of context and improves accuracy.
11. Be Careful With Connectors and Tools
Modern AI assistants can connect to an enormous amount of information (Google Drive, Slack, GitHub, web search, MCP servers). That’s powerful, but it’s also something you should use deliberately.
If you are asking an AI to rewrite a paragraph, you don’t need it searching the web, looking through your Drive, and checking GitHub. Turn off tools that aren’t needed for the task. Less capability can sometimes mean a more efficient workflow.
12. Create a Small Context File for Large Projects
This is one of the most useful habits if you regularly work with AI for coding or long-running projects. Instead of repeatedly explaining the same things, create a small document containing the project’s essential context.
Markdown
PROJECT.mdProject: Personal finance dashboardStack: Next.js, PostgreSQL, TailwindArchitecture: ...Coding conventions: ...Known issues: ...
Now a new AI conversation doesn’t have to reconstruct your project from scratch. The important thing is to keep this context small. Don’t turn your context file into another giant knowledge dump.
13. Know the Difference Between a Usage Limit and a Context Limit
This distinction is important enough to deserve its own section.
- A usage limit is about how much you can use the service over a period of time.
- A context limit is about how much information can fit into a particular conversation or request.
This explains why starting a new chat can sometimes help with a long conversation without magically giving you more overall usage. You’re reducing the amount of old context being carried forward, but you’re not bypassing the provider’s overall allowance.
The Takeaway: Better AI Hygiene
You cannot make a paid plan genuinely unlimited by cleverly formatting your prompts. If you genuinely use AI heavily every day, buying additional capacity is the right answer. But before doing that, look at how you’re spending the capacity you already have.
An enormous conversation containing five unrelated tasks is inefficient. Uploading a 100-page document to answer one question is inefficient. Using the most powerful model for grammar corrections is inefficient.
The goal isn’t to use AI less. It’s to waste less of the AI you already have. Keep conversations focused, give the AI the minimum context it actually needs, and ask for targeted changes instead of complete rewrites. Once you start thinking about AI this way, usage limits become much less mysterious. You aren’t trying to find a magic trick; you’re simply building a better workflow.
