Skip to content

claude-code

cc-limit-pacer drops Claude Code a model tier before lockout

Claude News

The number to remember from cc-limit-pacer is 41. In the sample month its author replays, only 41 of 2,410 Claude Code sessions ever grew past 500k tokens, so compacting earlier barely touches lockouts. The plugin, documented in its README on GitHub, targets volume instead: when your 5-hour or weekly window runs hot, it starts new sessions one model tier down and holds batch runs.

At a glance

  • cc-limit-pacer is a community Claude Code plugin, formerly called autocompact-gate, that watches your 5-hour and weekly usage through two hooks and changes settings only after you calibrate it and your usage runs hot.
  • In the README's sample replay of a month, the pacer plus compaction at about 830k tokens cut compactions by 28% and roughly halved the hours locked out, while $62 of batch work waited.
  • The replay leaves out quality loss from compacting or from cheaper models, and it also leaves out held work running later, so the author calls the lockout gains optimistic.

If you haven't been following, Claude Code plans meter usage in two windows, a 5-hour one and a weekly one, and /usage shows how much of each you've spent. Max out either and you're locked out until it resets. Fable also has its own weekly allowance, which only /usage reports. The plugin used to be called autocompact-gate, and on first run it moves its state from ~/.claude/state/autocompact-gate to ~/.claude/state/cc-limit-pacer.

In the sample month, the pacer cut locked hours from 74 to 38

The main evidence is a replay. The simulate command reruns your last month of transcripts under three policies. In the README's sample of 41,210 events over 4.3 weeks, the "today" row reproduces what actually happened: 18.0 compactions a week, 4 five-hour lockouts, 1 weekly lockout and 74.0 hours locked. The transcripts show 5 real five-hour lockouts and 1 weekly, and the README treats that close match as its sanity check.

Moving compaction to about 830k on its own gives 11.5 compactions a week. Lockouts stay at 4 and 1, though, and locked hours rise to 81.0, because the extra context brings lockouts sooner. The full install, 830k plus the pacer, comes out at 13.0 compactions a week, 2 five-hour lockouts, 1 weekly lockout and 38.0 hours locked, with $62 of batch work held back during hot stretches.

Only 41 of 2,410 sessions ever passed 500k tokens

The backtest command goes through the lockouts one at a time. In the sample it finds 6 lockouts across 2,410 sessions, with 71.0 hours locked in total. It then replays each one with early compaction while hot, at 500k and at 250k tokens. Most rows read "no help". In one Sep 19 lockout, for example, usage at the moment of lockout only falls from 103% to 95% of the limit.

The best case is a weekly lockout on Sep 20 that would have arrived 0.4 hours later. Across the whole set, the gate would have given back about 0.5 hours. Compacting every session at 500k would have cut usage by 3% over 32 days, because so few sessions ever got that far. The README's conclusion is that lockouts come from volume, which is why the other two levers exist.

The audit priced 70 auto-compactions at about $58

The audit command also reports what you wasted, over the last 23 days by default. In the sample data, 70 auto-compactions cost about $58 API-equivalent, or 1.3% of usage. The 4 session lockouts and 1 weekly lockout added up to 71 hours unable to work. Advisor consults came to $170, or 4% of usage, and Fable sat at 0% for the week, so its whole allowance went unused.

The stats command turns the same numbers into a page with meters, the savings table, the waste ledger, and weekly and 5-hour usage charts. A separate report command lists problems, such as a hook whose p95 latency is over 1 second, a calibration more than 3 days old, or a real lockout that happened while the pacer thought you were cool.

A window counts as hot at 50% used and on pace to run out

The pacer calls a window hot when it's at least 50% used and on pace to hit its limit. At that point it writes a model one tier down (Fable to Opus, Opus to Sonnet) and advisorModel: "off" into your settings, and shows a one-line notice. The change only reaches new sessions, so a session that's already running keeps its model. According to the README, "off" disables the advisor, but null and an empty string don't.

Automated runs, meaning SDK calls, claude -p and sessions working in a temp directory, are refused until the window resets, unless you set LIMIT_PACER_ALLOW=1. Interactive sessions are never held. It works like a household watching its power bill climb mid-month: the dishwasher waits for the cheap night rate and the heating goes down a notch, but nobody stops cooking dinner.

When usage cools down, the pacer restores your model and advisor settings, unless you changed them yourself in the meantime. The two hooks, SessionStart and UserPromptSubmit, take about 0.2 seconds each. Any exception goes to errors.jsonl and the prompt still goes through.

Calibration turns one /usage reading into a dollar budget

You start by passing your 5-hour percentage, your weekly percentage and the weekly reset time from /usage. In the sample, that produced budgets of 150 and 1,100 API-equivalent dollars for the two windows. After that, the hook estimates usage from local transcripts. Usage on claude.ai or on other machines doesn't show up, so the README tells you to recalibrate when status drifts away from /usage.

Plugins can't change autoCompactWindow, so a setup command backs up settings.json and sets it to 862000. Claude Code compacts about 32k tokens below that value, which puts the trigger at about 830k instead of about 570k on 1M-window models. If a single turn jumps past the model's real window, the session ends with "Prompt is too long" and nothing recovers it.

The README admits the replay is optimistic, since it ignores quality loss from compaction and cheaper models, and held work still has to run at some point. The usage sensor is approximate too, and it can't see other machines. In our view, the riskiest design choice is in headless mode: a held claude -p run returns the hold message as its result and exits 0, so a batch script that only checks exit codes will likely record a job that never ran as a success.

Getting into the private repo

The repo is private. Claude Code clones it with your normal git credentials, so git clone has to work for your account before the marketplace install does, and you need Python 3.9+ on your PATH. The README doesn't say whether the repo will become public. After you install, the number to watch is lockouts per week before and since install, which the status command shows. That comparison only tells you something once a few weeks have passed.

Related stories

  1. Claude Code gets a wrap-up budget at the 5-hour limit
  2. Anthropic deleted the post that called a cut a 25% raise
  3. Max20x buys only 1.5x of Max5x, says one team's tally
  4. A turn at 141 costs 2.1x a turn at 20 in Claude Code
  5. A Claude Code mod that reverts its own failed ideas
  6. Claude Design can draw with your real components

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.