Longer prompts cost more and answer worse; the fix is treating the AI traffic boundary as a control layer that trims context.