The following rant is not against the owner/project - but...
What an irony. I cant publish a attack surface mapping / pentesting tool i wrote which runs fully deterministic and really controlable due to "dual use" legal problems - but llm driven tools hit public space......
No im referring to the legal terms of germany, the country im residing at. Our laws regarding "hacking" are arguable the strictest and worst.
The problem is that they are formulated in a way that it is super easy to have your software being possible "dual use" and that a judge has to decide if its fine or not. Making it worse it also states your "intention" which well is impossible to proof - if the judge says he doesn't believe your intentions are only good, you can literally get massively sued.
So ye i could move to another country and than publish it - apart from that i can let it rot on my hdd (which is prolly what will happen).
Edit: Additionally mentioned, it is not just the publishing in germany, even the facilitating already which is why i don't even have an article about it (any more).
Completely understand, the legal landscape has really shifted around AI/LLM tools. I see tools drop everyday that spit in the face of DMCA/Copyright law but they skirt by mainly because they leverage AI
Agreed. Also less noticeable when "accidentally" dropped/left, especially in tight spaces out of view. And better odds of plausible deniability if caught or device is later found and somehow traced to you.
I built Nightcrawler, an open-source autonomous penetration-testing agent that runs entirely on an Android phone.
The project started with a question: how much of a real pentesting workflow could I run locally on relatively old mobile hardware, without relying on a cloud model or API?
Nightcrawler runs a 1.2B-parameter model locally on the Adreno GPU of a OnePlus 8. The model chooses targets and tools, while a separate scope-enforcement proxy validates every command before execution. The system maintains per-host memory in SQLite, rotates between targets, matches detected versions against a local CVE database, executes multi-step playbooks, and generates a structured report.
A few implementation details that may be interesting:
Local inference runs at roughly 115 prompt tokens/sec and 13 generated tokens/sec.
The small model only produces a usable command around 50% of the time, so much of the engineering is recovery logic, duplicate detection, persistent memory, and deterministic playbooks.
Every command passes through a separate scope and safety layer rather than trusting the model to remain in scope.
The project includes a dry-run mode, so the agent loop can be tested without executing real network commands or owning the phone hardware.
I've had it running on my home network for the past 3 months uninterrupted
So far only against four authorized networks. One of them was a corporate network. Left it overnight and it only found one minor week old CVE that I'm sure the IT team already had on their tracker. But showed that the proof of concept worked.
What an irony. I cant publish a attack surface mapping / pentesting tool i wrote which runs fully deterministic and really controlable due to "dual use" legal problems - but llm driven tools hit public space......
sorry for the rant....
Metasploit is one example: https://github.com/rapid7/metasploit-framework
The problem is that they are formulated in a way that it is super easy to have your software being possible "dual use" and that a judge has to decide if its fine or not. Making it worse it also states your "intention" which well is impossible to proof - if the judge says he doesn't believe your intentions are only good, you can literally get massively sued.
So ye i could move to another country and than publish it - apart from that i can let it rot on my hdd (which is prolly what will happen).
Edit: Additionally mentioned, it is not just the publishing in germany, even the facilitating already which is why i don't even have an article about it (any more).
The project started with a question: how much of a real pentesting workflow could I run locally on relatively old mobile hardware, without relying on a cloud model or API?
Nightcrawler runs a 1.2B-parameter model locally on the Adreno GPU of a OnePlus 8. The model chooses targets and tools, while a separate scope-enforcement proxy validates every command before execution. The system maintains per-host memory in SQLite, rotates between targets, matches detected versions against a local CVE database, executes multi-step playbooks, and generates a structured report.
A few implementation details that may be interesting:
Local inference runs at roughly 115 prompt tokens/sec and 13 generated tokens/sec. The small model only produces a usable command around 50% of the time, so much of the engineering is recovery logic, duplicate detection, persistent memory, and deterministic playbooks. Every command passes through a separate scope and safety layer rather than trusting the model to remain in scope. The project includes a dry-run mode, so the agent loop can be tested without executing real network commands or owning the phone hardware. I've had it running on my home network for the past 3 months uninterrupted