Inside the World of AI API Relay Services

In the spring of 2026, the somewhat gray-area business of AI relay stations suddenly welcomed a player who had no shortage of traffic.

The Most Opaque Infrastructure of the AI Era.

👦🏻 Author: GaKi

🧑‍🎨 Layout: NCon

In the spring of 2026, AI proxy services — a somewhat gray business — suddenly welcomed a player with no shortage of traffic.

In April, Justin Sun began promoting B.AI, calling the business of "one key to access all major models" the "underlying financial infrastructure for AI Agents."

This Chinese businessman, the crypto world's most gifted storyteller, nicknamed "Sun the Reaper" (reaper as in reaping韭菜, or retail investors), had suddenly fallen in love with the AI proxy business.

What he saw was obscene profitability. Pure obscene profitability.

But in reality, Justin Sun was merely a latecomer. Long before capital and industry heavyweights entered the game, a gray-market chain revolving around account arbitrage, splitting, and distribution had already been growing wild for quite some time.

And so an enormous, even ubiquitous "AI proxy underworld" was born.

🚥

What follows is our mapping and analysis of this "proxy underworld."

Proxies Are a Business with "Extremely Low Technical Barriers"

First, proxies have almost no technical barriers to entry, which has attracted hordes of "entrepreneurs" up and down the entire supply chain.

Under normal circumstances, requests go directly to Anthropic's or OpenAI's servers. With a proxy, you simply swap the request URL for its domain — it verifies your account balance, forwards to its upstream connection, and returns the result as-is.

You don't even need to build this architecture yourself. Deploy open-source projects like One API or New API, tweak a few configs, and you can have a proxy connecting multiple models, distributing keys, and calculating bills — all in a day.

So proxies online look virtually identical: unified API, multiple models, RMB top-ups, discounts starting at X% off.

But beneath the surface, they're actually four completely different businesses:

[1] Authorized resellers: Official procurement, then sold to you. Fully legal and compliant.

[2] Enterprise AI gateways: Companies pay for models themselves; the proxy only handles internal token allocation.

[3] Regional arbitrageurs: Overseas purchasing, then resold to domestic users.

[4] Gray-market proxies: Using违规手段 (subscription pools, fake accounts, model swapping) to drive costs extremely low.

The same /v1/chat/completions endpoint, the same Claude Code access — but cost structures and shelf lives are completely different.

When discussing proxies, one can't avoid mentioning OpenRouter, the benchmark in aggregation. While OpenRouter doesn't fit the typical proxy model we'll focus on — one that operates through "gray arbitrage, account pools, and clandestine distribution" — as the internet's largest legitimate aggregation platform, it provides a useful reference point.

OpenRouter makes every provider's and every model's pricing and data policies fully public, charging only a platform fee on top-ups.

But many don't use it, because while it offers genuine Claude API access, the prices are simply too high — essentially official rates plus roughly 5.5% in fees. So the industry joke goes: calling Claude through OpenRouter is more expensive than calling the official API directly.

Meanwhile, domestic low-price proxies advertise rates at roughly one-tenth of Claude's official price.

How Do Proxies Actually Make Money?

The most common domestic proxy ads read "1 RMB tops up 1 USD" or "Everything 20% off." This works because several profit-hidden links never make it onto the price list.

The "USD" in your balance isn't dollars — it's the proxy's own point system. Recharge exchange rates and model multipliers are all set by the site operator.

So the same displayed 100 USD can buy vastly different amounts of tokens across sites: some give 5 USD per 1 RMB charged at 1x, others give 1 USD per 1 RMB at 0.2x — the actual cost may end up identical.

We had Codex map out the math, and here's roughly what it "told" us the payment logic looks like (R = USD received per 1 RMB, M = combined billing multiplier for model/user group/routing, L = official USD list price):

Credits received: B = Actual RMB paid A × R
USD deducted: C = Official equivalent price L × Combined multiplier M
Actual RMB cost: Y = C ÷ R = L × M ÷ R
RMB cost per 1 USD of official list price: E = M ÷ R

Complex enough? Just looking at recharge ratios, you can't tell who's cheaper. And that confusion is precisely what creates markup room for proxies. This same logic appears on Justin Sun's B.AI.

Where Do Profits Actually Come From?

Today, most proxies' main service isn't aggregation at all — most only offer Claude and GPT 5.6 series, with domestic models like DeepSeek even thrown in as free perks.

To turn a profit, model sourcing can't possibly be through normal channels, let alone official Claude API.

Profit pathways roughly break down as follows:

[1] Wholesale discounts.

Securing below-retail prices through enterprise contracts or bulk purchasing, then selling near official rates — pocketing the spread. The cleanest path, but thinnest margins. Can't sustainably drive flagship models to extreme lows this way.

[2] Subscription splitting (the primary path).

Operators buy large quantities of Claude and ChatGPT subscription accounts — say, Claude Max 5X. Purchased in bulk, sources are mostly gray. A glance at Taobao account sales gives the full picture.

Claude strictly prohibits account sales, sharing, and purchases from unsupported regions — ban probability is extremely high. Yet sellers offer warranties: banned after 7 days? Minus fees and usage for those 7 days, full refund of the rest.

This warranty itself proves the account sources are deeply gray.

Operators then pool these accounts, converting actual subscription usage into a much larger nominal credit pool at official API unit prices.

Corresponding model multipliers adjust with ban probability — prices spike dramatically during the strictest crackdown periods.

So this isn't stable margin — it's closer to statistical gambling.

Pool size, ban probability needed to sustain user volume, account lifespan, concurrency caps, whether quotas get fully used — all must be calculated in advance. One rule change or mass ban from upstream, and profits zero out.

[3] Downstream distribution.

Large sites give agents lower wholesale rates; agents handle user acquisition, community management, and customer service, marking up layer by layer to end users. The site becomes a wholesaler rather than retailer.

[4] Silent model substitution.

Early proxies still advertised having dozens of brands; now they target only Claude and GPT. Meanwhile, domestic models have improved rapidly — the previously popular GLM 5.1 was heavily used to replace Claude's API. It's extremely cheap, and domestic oversight of such reverse-proxy subscriptions is relatively lax.

The price gap is enormous. Packaging cheap models as expensive ones — profit comes directly from users unable to tell real from fake.

However, with Claude Pro finished accounts now selling for rock-bottom prices on the secondary market, basically approaching "official Claude at original price," many operators have lost the incentive to "secretly substitute" fake models for genuine Claude to retain users.

So, How Much Can a Proxy Actually Make?

First, look at the most legitimate: OpenRouter. Covering 70+ providers, 400+ models, 10 million users, processing 200 trillion tokens monthly — no cheap models, no account arbitrage, yet substantial scale.

As of March 2026, annualized revenue exceeded $50 million. In May, it closed a $113 million Series B at roughly $1.3 billion valuation. On July 23, The Wall Street Journal reported Stripe was in acquisition talks, with potential deal value reaching approximately $10 billion.

Even fully compliant, just doing API aggregation — that's enough to support a unicorn.

Now look domestically.

How much can these proxies actually earn? Publicly reported figures already span from grassroots one-person sites to platforms at the tens of millions level.

The smallest category might be personal sites built by individuals. China Newsweek, citing previous media reports, mentioned an operator who previously worked in a factory workshop, barely understood programming, simply followed tutorials and had AI help build the website, then slapped on the New API open-source framework — and eventually got real paying users.

This site averaged thousands to over ten thousand RMB in daily top-ups. Rough back-of-envelope: monthly top-up flow somewhere between 30,000 to 300,000 RMB.

Scaling up, public reports have already surfaced projects with monthly flows in the millions. In May 2026, 21st Century Business Herald (reprinting a Jiemian News investigation) reported that an AI proxy project tracked by one investor had approximately 5 million RMB monthly flow with gross margins near 50%.

In a June 2026 Baobian report, a relatively new proxy operator claimed that "some companies in the industry doing monthly flows at the tens of millions level" operate with teams of under 20 people.

This reveals an exaggerated industry picture: if a sub-20-person team hits 10 million RMB monthly top-up flow, that's over 500,000 RMB monthly flow per employee. Annualized, platform top-up flow could exceed 100 million RMB.

The key behind high flow: proxies can drive actual upstream model call costs far below official API prices. Baobian cited one typical cost model: assuming official API prices 1 RMB per 10,000 tokens, proxies can dilute costs to roughly 0.2 RMB through bulk Coding Plan purchases, package merging, and account pooling — then sell at 0.5 RMB. Under this model, every 0.5 RMB of tokens sold costs ~0.2 RMB, yielding 0.3 RMB gross profit, or roughly 60% gross margin.

Low costs ultimately reflect in external pricing. Baobian cited a 21st Century Business Herald investigation finding that Claude Opus 4.6's official API output price was approximately 170 RMB per million tokens; one domestic proxy had marked this down to roughly 50% of official price, with smaller sites at just 20-30%.

But proxies aren't guaranteed money. Sites launch for a few thousand RMB; real costs lie in account pools, proxy IPs, payment channels, ban losses, refund payouts, and replenishing upstream resources. The result is polarization: leaders capture high margins through scale and channels, while small and mid-sized sites hover at break-even or loss.

On May 12, 2026, social media account @YesterdayBigcat reposted a statement screenshot signed "Watermelon Skin": the poster identified as a Shanghai proxy operator, detained for 37 days on criminal charges for illegally obtaining and reselling API resources (some low-price APIs obtained through "illegal technical means"), released on bail after completing restitution, and facing fines.

He also stated the proxy ultimately didn't make money, even lacking funds to refund remaining user balances.

High-profit cases do exist, but those who actually make money tend to be just the handful at the very top of the chain.

Data Reselling

Beyond account and model-level tricks, what's easier to overlook is data.

All requests passing through Coding Agent go via the proxy's servers — complete code repository context, prompts, error logs, even keys leaked during debugging are theoretically visible to it.

Whether this gets retained or used for other purposes, users have no way of knowing, no way of verifying.

In May 2026, Oxford China Policy Lab researcher Zilan Qian, investigating China's gray Claude proxy market, interviewed multiple developers. Their take: API markup is just customer acquisition; collecting call logs is the real profit source.

These logs contain user inputs, model outputs, error messages, and human-corrected results — usable for model distillation, code model training, product analytics, or resale.

Per Bright Data, ordinary web or structured records on open markets run as low as tens to hundreds of USD per 100,000 entries.

Of course, scraped web data and high-quality reasoning traces aren't the same commodity. And high-quality data under Claude Opus 4.8 versus Fable 5 models command different prices entirely.

Wherever AI Can Be Used, Proxies Have Already Infiltrated

Proxy use cases have grown so widespread that many don't even realize they're using one.

The most classic公开 (public-facing) use: buy API, get a key, plug into Codex, Claude Code.

(CC Switch freely connecting to Claude proxy API for use in Claude Code)

But dig deeper, and commercial proxy scenarios get far more varied — and already exceed most people's "trust comfort zone." For instance, quietly embedded into various AI coding product backends, replacing original AI requests, secretly rerouting to their own API proxy.

This mainly happens because the vast majority of users have no way to verify which model they're actually getting. Proxies need only pre-write a system prompt — no matter how users try to trick or prompt it into revealing its identity, they basically can't get the truth.

Any product using AI capabilities is theoretically susceptible to such substitution.

There's a particularly classic case here.

Someone online sells a "Cursor unlimited Claude membership" solution — absurdly cheap, thousands of calls daily. The package includes a tutorial teaching you to install a plugin in Cursor; the plugin displays an account, deducting per call based on package, even supporting one-click account switching.

(This image shows the Cursor interface, with the "plugin" on the left)

The tutorial specifically emphasizes manually switching to natively supported Claude models in Cursor's right-side menu. Users see they've manually switched to Cursor's built-in model and subconsciously assume this must be real.

After all, the plugin's call count does decrease authentically — it looks legitimate.

Until Codex checks the backend code and discovers it never touched Cursor's native model logic at all — instead secretly redirecting the model source address to a website that, when opened, is an API proxy.

Accounts, call counts, one-click switching, even the "manually switch models" guidance itself — all smokescreen.

Whether the actual calls are really Claude is completely unclear — after all, thousands of calls for a few dozen RMB, the price itself already says plenty.

So proxies enter products through far more than one entry point; they're dispersed across every corner using AI capabilities, and users often don't even know they've been hijacked.

This also has third-party confirmation. A CISPA research team spot-checked 17 proxy interfaces cited in over a hundred academic papers; some performance deviations reached 47%, with nearly half failing model fingerprint tests.

One interface labeled Gemini 2.5 scored just 37 on medical benchmark tests — official interfaces score above 83.

Price alone can't determine whether a proxy's model is genuine.

And such scoring has itself become a market: now there are numerous API proxy review sites responsible for giving each site's models "match scores" — yet this scoring itself takes only 10-20 seconds, so whether the scores are real, users again have no way of knowing.

Why Does Such a Complex Proxy Underworld Exist?

The root cause: official channels simply aren't available domestically — OpenAI and Anthropic's supported regions lists exclude China, legitimate registration requires overseas phone numbers, payment requires overseas credit cards, and KYC grows stricter year by year.

Especially Anthropic, recently exposed for using bait emails to identify regions, and even planting various tricks in Claude Code to determine if users are from China.

So for ordinary developers facing these frontier AI tools, the first problem is simply whether they can use them at all. Proxies solve "can use" before "cheap" — this was the business's original foothold.

Later, Claude Code, Codex and other coding agents amplified demand. Previously, using models meant asking one question, getting one answer — occasional emergency top-ups were no problem.

But automated workflows and long-context workflows require SOTA-level agentic models to extensively read files, call tools, compress context — token consumption is continuous. The same task via API might cost 5x, 10x, even 20x what a subscription membership would.

Those who can't access subscription memberships are easily lured by cheap advertising — and fall into the trap. In the era of coding agents, proxies have essentially become daily infrastructure that many developers can't do without.

Growing demand complexified supply chains. The entire proxy underworld roughly divides into three layers:

[1] Accounts and resources: Official API keys, enterprise discounts, account pools and reverse-proxy tech, even overseas stolen credit cards and SMS verification services;

[2] Proxy operators: Managing accounts, routing, and remediation after bans;

[3] Users and各级代理 (multi-level agents): From second-tier to fourth-tier, individual developers, enterprise teams, and small agents who wholesale credits from large sites for resale.

The biggest characteristic is modularity. One site gets banned — account sellers, account pools, card sellers, technical solutions are all ready-made, almost zero migration cost, just swap domains and reopen.

This is also why this business operates like an "underworld," not like a company.


In a sense, proxies are the most contradictory class of infrastructure in the AI era.

On one hand, they lower barriers to using advanced models, letting many who otherwise couldn't access Claude, GPT, and coding agents get started first. On the other, they compress trust, compliance, data security, and model authenticity into a black box.

They're used as frequently as water, electricity, and gas — yet unlike true infrastructure, they're auditable, accountable, and reliably long-term.

This is why proxies won't remain mere footnotes in gray-market news. They actually expose a larger problem: as AI capabilities increasingly resemble means of production, whoever can stably, cheaply, and credibly distribute these capabilities becomes a new node of power.

Today that node might be OpenRouter, might be B.AI, might be some small operator hiding in Telegram groups and WeChat groups.

The only difference: some built it as a company, others built it as an underworld.

Crossing is seeking independent contributors to write AI product and model reviews.

If you've written articles like: "[Hands-on] PixVerse C1[1]", "[Hands-on] LibTV[2]", please contact zeo0811@gmail.com. Email should include: ① personal introduction, ② AI review articles you've written.

We offer competitive rates. Looking forward to observing and documenting the AI era with you 🎪

References

[1] [Hands-on] PixVerse C1: https://mp.weixin.qq.com/s/cgAzZy2PptdYaWbqywVg7A

[2] [Hands-on] LibTV: https://mp.weixin.qq.com/s/aycqnq8wlaaOei1QZmhFHQ