Scanner Collect: Data Ingestion
Learn what Scanner Collect is, how it delivers your SaaS and cloud logs into S3 storage you control, and how it simplifies ingestion, indexing, and detection.
Scanner Collect delivers logs from all your sources into Amazon S3 with minimal setup and no custom pipeline code. In a single afternoon, you can integrate dozens of log sources.
Logs are delivered into your S3 buckets as gzipped JSON files, partitioned by date (e.g. s3://your-bucket/logs/okta/2025/07/27/). Scanner can then optionally index these logs to enable full-text search and continuous detections.
Why Land Logs in S3?
Most traditional logging tools and SIEMs rely on stateful search clusters, which are costly and difficult to scale. Landing logs in S3 object storage offers key advantages:
Orders-of-magnitude lower storage cost
Easy to scale from GBs to PBs
Native support for long-term retention
Your logs stay in your storage and under your control
Scanner Collect simplifies the ingestion layer, handling API polling, file formatting, delivery to S3, and accelerated full-text search and detections.
The Modern Security Perimeter Is SaaS and Cloud
Security monitoring today means more than just collecting logs from endpoints or firewalls. The modern attack surface is defined by SaaS tools, identity providers, and cloud platforms, and each of them emits valuable audit logs.
Across a large organization, that log data accumulates in three kinds of places: object storage, SaaS tools, and data lakes or warehouses. Scanner Collect has an ingestion mode for each, so your team can connect log sources in minutes, not weeks, and focus on what matters: detection, investigation, and incident response.
Log Source Types
1. Object Storage
Logs that already sit in cloud object storage are indexed without moving your originals:
Amazon S3 - Indexed in place. Scanner receives
s3:ObjectCreatednotifications as new files land (JSON, Parquet, CSV, or plaintext) and indexes them for search and detection. This covers any logs already delivered to S3 by third-party pipelines or vendor tooling: Cloudflare DNS and HTTP logs, Crowdstrike Falcon Data Replicator, Sublime Security email logs, and more. See Create an Index Rule.Google Cloud Storage and Azure Blob Storage - A mirror pipeline copies new objects into a short-lived S3 collect buffer, where Scanner indexes them like any other S3 source; your original objects are never modified, moved, or deleted. Objects move gzip-compressed and expire from the buffer after 7 days by default, so cross-cloud transfer and storage costs stay low while Scanner's index retains the searchable data. See Google Cloud Storage (GCS) and Azure Blob Storage.
2. SaaS Tools
Scanner Collect connects to SaaS sources in one of two ways, depending on how the tool makes logs available:
API Pull
Scanner integrates with API-based log sources and periodically pulls logs into S3. It handles API pagination, authentication, and deduplication internally.
Examples:
Okta System Logs (via
/api/v1/logs)Google Workspace Admin Activity Reports (via Reports API)
Slack Audit Logs (via
auditlogs.slack.comAPI)
Behavior:
Logs are fetched periodically (typically every 1–5 minutes)
Delivered to S3 as
.json.gzfiles under a daily partitioned prefixEach file is a newline-delimited JSON file, where each line is a new log event
HTTP Push
Scanner can accept logs over HTTP, useful for tools that emit webhook events or for forwarding from log shippers.
Examples:
Alert webhooks from Wiz, Google Alert Center, Tines, Torq
Logs pushed from Fluent Bit, Logstash, or Vector
Behavior:
Scanner exposes a secure HTTP endpoint per source
Incoming payloads are batched and written to S3 in gzip-compressed JSON
Custom parsing/enrichment pipelines can be configured as needed
3. Lakes & Warehouses
Many teams land the same events in Snowflake, Databricks, or ClickHouse through streaming infrastructure. Rather than querying the warehouse, Scanner taps the streams that feed them (Kafka, Kinesis, MSK, Cribl, and others): point the stream's S3 sink at a collect buffer bucket, and the same events become searchable and monitored in Scanner without touching the warehouse itself. See Streams (Kafka, Kinesis, and More).
Forward Elsewhere (Optional)
Scanner itself does not forward logs to other systems after ingestion. However, because your logs are stored in your own S3 bucket in a clean, consistent format, you’re free to set up your own forwarding pipelines.
Common approaches include:
Triggering AWS Lambda functions on new S3 object creation
Streaming from S3 to Kinesis, then into another SIEM (e.g., Splunk, Datadog)
Using AWS Glue or other ETL tools to load logs into downstream systems
This flexibility is intentional. You retain full ownership of your log data and can integrate it into any part of your security stack without vendor lock-in.
Last updated
Was this helpful?