80,000 Relay Servers Found Masking AI Model Queries
Researchers have uncovered a vast network of AI gateways that conceals the real users behind requests sent to Claude, OpenAI, Gemini, xAI, and other models. Team Cymru first counted 10,867 such servers, then, after broadening its search, revised the estimate to more than 80,000 relays. A significant portion of the traffic examined originated from China and Hong Kong, and certain clusters may serve not only to bypass regional restrictions but also to harvest the responses of frontier models en masse. Team Cymru laid out the scheme in its report on how LLM gateways enable frontier model abuse.
How the Gateways Work
Such servers function as an intermediary layer between the user and the AI provider. The client sends a request to the gateway, which selects an account, API key, or active user session from its pool and then contacts the model under its own identity. As a result, the AI provider sees the intermediary’s address and credentials, but not the end user’s real IP.
The architecture itself is not inherently criminal. Corporate AI gateways are used for centralized request accounting, load balancing, cost control, and working with multiple models at once. The trouble begins when a middleman pools other people’s accounts or accounts registered in bulk, resells access, hides the client’s geography, or helps circumvent a provider’s restrictions.
Claude Relay Service and Its Successor sub2api
Team Cymru identified two especially widespread open-source projects, Claude Relay Service and its successor, sub2api. The latter can manage users, set individual limits and billing, turn subscriptions into an API-like service, and audit requests. According to the researchers, sub2api has drawn more than 8,000 forks on GitHub, and its associated Telegram channel has gathered nearly 7,000 subscribers.
The researchers also examined the commercial ecosystem surrounding such gateways. Among sub2api’s sponsors, Team Cymru found sellers of access to model APIs, residential-proxy providers, account sellers, and infrastructure services geared toward heavy AI-request traffic. Sponsorship alone does not prove any particular company engages in illegal activity, but the range of services shows just how far the AI-relay market has drifted from ordinary corporate proxying.
Inside One Cluster: 14TB Up, 7TB Down in Eight Days
In one cluster examined, more than 4,000 IP addresses from China and Hong Kong connected to 304 intermediary servers. Over eight days, clients transferred roughly 14TB of data to the relays and received more than 7TB back. After excluding minor flows, about 244 active addresses remained, which reached out to 173 gateways, which in turn connected to 262 AI services and other related resources.
One group of gateways addressed the Chinese models DeepSeek, Qwen, Doubao, Zhipu, and MiniMax. Team Cymru allows that such activity could simply be ordinary aggregation and resale of access. Another group connected to Anthropic, OpenAI, Google, and xAI. It was this second cluster that drew greater attention, owing to the geography of its users and an unusual traffic distribution.
A Suspicious 58-to-1 Traffic Ratio
Through 17 relays reaching api.anthropic.com, about 81GB of data went out over eight days, while roughly 1.4GB came back in. The ratio of outbound to inbound traffic came to about 58 to 1. Team Cymru regards so pronounced a skew as a potential sign of mass request preparation for model distillation, but explicitly stresses that the researchers never saw the content of the requests. The observed activity therefore cannot yet be called a proven attack on Anthropic.
Distillation itself is an ordinary AI-training method. A more powerful model acts as a teacher, generating answers, while a smaller model is trained to reproduce part of the capability gained. Developers apply this approach legitimately to build cheaper, more compact models. The violation arises when an outside company harvests a competitor’s responses en masse, against the terms of service, and uses them to train its own system.
Anthropic’s Own Findings
Team Cymru’s concerns did not arise in a vacuum. In February, Anthropic reported detecting campaigns by DeepSeek, Moonshot, and MiniMax that, by the company’s account, created roughly 24,000 fraudulent accounts and conducted more than 16 million interactions with Claude. In a September report, Anthropic described fresh campaigns from seven Chinese labs and states that some operations generated millions of requests per day. These findings belong to Anthropic itself and do not prove that the relays Team Cymru found took part in those same operations.
Regional Restrictions Create Another Market
Regional restrictions give the gateways yet another market. OpenAI warns that accessing its API from countries and territories outside its list of supported regions can lead to an account being blocked or suspended. China is not on that published list. Anthropic likewise provides no commercial access to Claude in China and links part of the discovered proxy networks to attempts to circumvent such restrictions.
A Deeper Problem Than Simple IP Rotation
For AI providers, a large-scale relay network creates a problem deeper than a simple change of IP address. Defense systems tie limits, billing, geography, reputation, and signs of abuse to a specific account and connection point. A gateway severs that link, mixing the requests of different clients behind a single set of credentials. Once one account is blocked, the operator can simply switch the flow to the next, so combating this kind of infrastructure begins to resemble the fight against ordinary proxy networks and botnets.
Team Cymru has passed the discovered addresses on to AI companies and continues to search for new nodes. The study’s central limitation remains unchanged. Network analysis reveals routes, volumes, and participants in connections quite well, but it does not automatically expose each user’s intent. The proven result, for now, remains the existence of a mass infrastructure that hides the origin of AI requests and allows part of the control mechanisms to be bypassed. The use of specific nodes to steal a model’s capabilities is something the researchers can, for now, establish only in isolated cases.
Support Our Threat Intelligence
If you find our technology report and cybersecurity news helpful, consider supporting our work.