Active lighting tests that preserve color-dependent capability
Robot QA teams should include spotlight hue, intensity, position, and beam angle as factors in active real-world evaluation, while separating tasks that require color from those that do not. FLARE shows why failure discovery alone is insufficient: broad color augmentation survived attacks partly by teaching policies to ignore color, cutting benign success on a real color-dependent task to 47.5%. A factor-based surrogate could instead select informative lighting configurations and track both attack robustness and retained color discrimination, reducing physical trials without rewarding this shortcut.
The cheapest check is to add grayscale and color-swap diagnostics to an existing active-testing run, then compare the discovered failure regions with uniform-random lighting tests under the same hardware budget.