OpenAI has published a practical guide to its GPT-6 family, positioning the suite as a set of models for different kinds of work rather than a single do-everything system. The guide walks through three tiers: GPT-6 Astra for the hardest reasoning, GPT-6.1 Sol for complex coding and research, and GPT-6 Luna for high-volume, well-defined tasks like extracting invoice fields or classifying requests. It also advises setting reasoning effort explicitly, from low for routine extraction to high or extra-high for difficult debugging and review.

Cost and latency management is a central theme. The guide recommends cutting unneeded context, running independent tasks in parallel, and using prompt caching, noting that cached input tokens can cost up to 95% less than uncached ones. For long conversations, compaction reduces context size while preserving state. OpenAI also urges teams to measure task success, latency, and cost per successful task before deploying, and to review monitoring and data controls.

On prompting, the guide argues that overly specific instructions can now hurt results because models handle nuance better. It suggests giving a clear assignment with a definition of done, keeping skill descriptions short, updating AGENTS.md files to authorize safe workflows, and setting explicit decision boundaries for when the model should act independently versus ask for approval. The advice is practical and product-oriented, with no independent evaluation of the claims.