Rate Limiting Strategies and Implementation
Rate Limiting Algorithms
The platform supports three rate limiting algorithms configurable per endpoint: fixed window, sliding window log, and token bucket. Fixed window counts requests in discrete time intervals (e.g., 100 requests per minute), which is simple but susceptible to burst traffic at window boundaries. Sliding window log maintains a sorted set of request timestamps and counts requests within a trailing window, providing smoother rate enforcement at higher memory cost.
Token bucket is the default algorithm and provides the best balance of fairness and burst tolerance. Each API key has a bucket that fills at a constant rate (tokens per second) up to a maximum capacity (burst limit). Each request consumes one token. When the bucket is empty, requests are rejected until new tokens accumulate. Configure per-endpoint: PUT /org/settings/rate-limits with {"endpoint": "/scans", "algorithm": "token_bucket", "rate": 10, "burst": 50}.
Custom Rate Keys
By default, rate limits are applied per API key. For more granular control, configure custom rate keys that combine multiple dimensions. For example, rate limit by API key and endpoint: different endpoints can have different limits for the same key. Or rate limit by API key and source IP: prevents a single compromised client from consuming an organization's entire quota.
Configure custom rate keys via PUT /org/settings/rate-limits/keys with {"dimensions": ["api_key", "endpoint", "source_ip"], "combination": "all"}. The combination field specifies whether limits apply to all dimensions combined (all) or to each dimension independently (any).
Graceful Degradation
When rate limited, the API returns a 429 Too Many Requests response with Retry-After and X-RateLimit-Reset headers. Implement exponential backoff with jitter in your client: wait min(2^attempt * 1000 + random(0, 1000), 30000) milliseconds between retries. The official SDKs handle this automatically.
For high-availability integrations, configure a webhook fallback: when rate limited on real-time API calls, the platform queues events and delivers them via webhook when capacity is available. Enable fallback via PUT /org/settings/rate-limits/fallback with {"webhook_url": "https://yourapp.com/hooks/fallback", "buffer_ttl_seconds": 3600}. Buffered events are delivered in order with at-least-once semantics.
This article is part of our ongoing security research series. Related data is available through the linked endpoints.