Technology
Danish Kapoor
Danish Kapoor

Microsoft released a tool that measures errors of artificial intelligence agents

In the workflow described by Microsoft, the developer first describes the agent in natural language. The component called Clarity extracts possible error paths; ASSERT prepares evaluation scenarios for these. Then, the rule to be applied during operation is created with the Agent Control Specification. In the final stage, the agent is tested again with the same behavioral definition, the same tests and the same evaluation approach. Thus, the possibility that the difference between the first and second results is due to changes in test conditions is reduced.

The company’s example is based on a billing agent that only needs to serve one customer. In the first measurement, the agent exposed another customer’s data in 12 of the 40 relevant conversations; the rate is 30 percent. In the example after adding the rule that compares the account ID with the user’s authority, two of the 34 relevant conversations were violated; The rate decreased to 5.9 percent. The figures are the result of this case study. It does not show that the same improvement will be achieved in every agent or real company environment.

At what point does the security rule come into play?

In the example, the rule runs before the agent calls the agent requesting data from another account. Inappropriate call is blocked before execution; Then, a second check is applied to ensure that an erroneous result does not return to the agent’s context. This approach differs from simply reviewing the response text later. It draws boundaries when it comes to accessing sensitive data. In addition, the developer needs to review the rule and the point at which it is applied; The generated policy is not automatically approved for distribution.

Microsoft says it measures secure behavior and ability to assist legitimate requests separately. Because an agent that rejects every request may not leak data, but it also cannot do its job. In the split test results shared by the company, while the rates of undesirable behavior decreased, violations of permitted behavior also decreased. However, in some test groups, security breaches were not completely eliminated; The tool should not be seen as a one-time definitive solution.

Run-assert-eval is available as open source in the ASSERT repository. The real benefit for the developer is to keep the risk assumptions, the rule, and the measurement of the outcome in the same loop. On the other hand, the percentages in the samples come from limited test sets. When different users, tools and permissions come into play in the production environment, new error paths may arise. Therefore, human review before release and re-evaluation in subsequent versions are still required.

Danish Kapoor