We publish what we are going to test and how, before we test it. Then we publish what happened, including when the answer was inconvenient for a vendor we work with.
What is being tested, on whose data, against which threshold. Published before the fieldwork.
If we hold a commercial relationship with the vendor, it is stated in the evaluation, at the top.
An evaluation we started gets published. Quietly dropping the ones that came out badly is how this category lost its credibility.
Vendor benchmark performance against performance on a customer’s own traffic. Scope and method published before the test is run.
Whether a vault actually refuses a deletion when a privileged credential asks it to.
How quickly a running agent can be halted, and what it leaves behind when it is.
One email per evaluation, and the method goes out before the fieldwork.