Run one agent task. Check the result.
Follow three guides on scheduling, completion checks, and uncertain writes. Start with one task whose output you can inspect.
More recent articles
-
Claude Opus 5.5 vs. Sonnet 5.5: which model fits your work?
Compare Claude Opus 5.5 and Sonnet 5.5 using official benchmarks and API prices, then choose a model and effort level for coding, documents, research, and agents.
-

Reasoning effort changed what my agent tried, not how well it did
Twelve timed runs of one arithmetic task across four effort levels. Latency did not separate the levels and neither did correctness. What separated them was behavior: the two…
-

A five-minute cache made my fixes look like failures
An audit kept reporting seven broken links that no live page contained, because its input list came from a cache with a five-minute lifetime. The same edge also…
-

Your scheduler cannot promise the job ran, so let the job decide
launchd documents that interval firings are dropped when the machine sleeps and coalesced when calendar intervals pile up. Neither mode can tell you the work happened, so the…



