A new benchmark named KaliBench has been developed to specifically evaluate the ability of large language models (LLMs) to translate natural language queries into precise, executable command-line interface (CLI) commands for cybersecurity tools. The benchmark addresses a gap in existing evaluations, which often focus on knowledge-based assessments rather than the direct generation of functional commands. Cybersecurity operations demand exact syntax in CLI commands, where even minor errors can prevent execution.
KaliBench comprises 8,504 query-command pairs, covering 1,642 tools across 23 capability dimensions and five security phases within Kali Linux. The dataset was constructed using a "manuscript-grounded pipeline" that includes deterministic canonicalization and alias-aware evaluation. To ensure the quality of the data, a multi-stage verification process was employed. This process combines LLM-based validation, sandboxed terminal execution, and human review to confirm both the semantic correctness and practical executability of the commands.
Initial testing with KaliBench involved 24 configurations of both general-purpose and security-focused open-weight LLMs. The results showed that none of these models exceeded 42% exact-command accuracy in the unrestricted evaluation setting. This outcome highlights the difficulty LLMs face in generating precise CLI commands for cybersecurity tools without explicit hints.
The researchers also demonstrated that supervised fine-tuning and reinforcement learning, utilizing the verifiable rewards provided by KaliBench, can significantly improve the performance of an 8-billion-parameter model. This improved model achieved performance comparable to a 685-billion-parameter Mixture-of-Experts (MoE) model. KaliBench's design allows for "runtime-free verifiable rewards" for training, which means models can be evaluated and receive feedback without needing to execute commands in a live environment for every assessment. This feature is built on the benchmark's deterministic canonicalization and alias-aware evaluation mechanisms.
The development of KaliBench provides a fine-grained method for assessing LLMs in a critical aspect of cybersecurity. The benchmark aims to facilitate the development of more capable LLMs for cybersecurity workflows, where they are increasingly expected to assist analysts by invoking tools based on their intent.
