Skip to main content
Version: 1.5.0

IBM Cloud Object Storage

IBM Cloud Long Term Storage

Synopsis​

Creates a target that writes log messages to IBM Cloud Object Storage buckets with support for various file formats and authentication methods. The target handles large file uploads efficiently with configurable rotation based on size or event count. IBM Cloud Object Storage provides enterprise-grade durability, security, and global availability.

Schema​

- name: <string>
description: <string>
type: ibmcos
pipelines: <pipeline[]>
status: <boolean>
properties:
key: <string>
secret: <string>
region: <string>
endpoint: <string>
part_size: <numeric>
bucket: <string>
buckets:
- bucket: <string>
name: <string>
format: <string>
compression: <string>
extension: <string>
schema: <string>
name: <string>
format: <string>
compression: <string>
extension: <string>
schema: <string>
max_size: <numeric>
batch_size: <numeric>
timeout: <numeric>
field_format: <string>
interval: <string|numeric>
cron: <string>
debug:
status: <boolean>
dont_send_logs: <boolean>

Configuration​

The following fields are used to define the target:

FieldRequiredDefaultDescription
nameYTarget name
descriptionN-Optional description
typeYMust be ibmcos
pipelinesN-Optional post-processor pipelines
statusNtrueEnable/disable the target

IBM Cloud Object Storage Credentials​

FieldRequiredDefaultDescription
keyY-IBM Cloud Object Storage HMAC access key ID
secretY-IBM Cloud Object Storage HMAC secret access key
regionY-IBM Cloud region (e.g., us-south, eu-gb, jp-tok)
endpointY-IBM COS endpoint URL (e.g., https://s3.us-south.cloud-object-storage.appdomain.cloud)

Connection​

FieldRequiredDefaultDescription
part_sizeN5Multipart upload part size in megabytes (minimum 5MB)
timeoutN30Connection timeout in seconds
field_formatN-Data normalization format. See applicable Normalization section

Files​

FieldRequiredDefaultDescription
bucketN*-Default IBM COS bucket name (used if buckets not specified)
bucketsN*-Array of bucket configurations for file distribution
buckets.bucketY-IBM COS bucket name
buckets.nameY-File name template
buckets.formatN"json"Output format: json, multijson, avro, parquet
buckets.compressionN-Compression algorithm. See Compression below
buckets.extensionNMatches formatFile extension override
buckets.schemaN*-Schema definition file path (required for Avro and Parquet formats)
nameN"vmetric.{{.Timestamp}}.{{.Extension}}"Default file name template when buckets not used
formatN"json"Default output format when buckets not used
compressionN-Default compression when buckets not used
extensionNMatches formatDefault file extension when buckets not used
schemaN-Default schema path when buckets not used
max_sizeN0Maximum file size in bytes before rotation
batch_sizeN100000Maximum number of messages per file

* = Either bucket or buckets must be specified. When using buckets, schema is conditionally required for Avro and Parquet formats.

note

When max_size is reached, the current file is uploaded to IBM COS and a new file is created. For unlimited file size, set the field to 0.

Scheduler​

FieldRequiredDefaultDescription
intervalNrealtimeExecution frequency. See Interval for details
cronN-Cron expression for scheduled execution. See Cron for details

Debug Options​

FieldRequiredDefaultDescription
debug.statusNfalseEnable debug logging
debug.dont_send_logsNfalseProcess logs but don't send to target (testing)

Details​

The IBM Cloud Object Storage target provides enterprise-grade cloud storage integration with comprehensive file format support. IBM COS offers high durability (99.999999999%), security features, and flexible storage classes for cost optimization.

Authentication​

Requires IBM Cloud Object Storage HMAC credentials. HMAC credentials can be created through the IBM Cloud console and provide programmatic access to COS buckets.

Endpoint Configuration​

IBM Cloud Object Storage uses region-specific endpoints. The endpoint format is typically https://s3.<region>.cloud-object-storage.appdomain.cloud where <region> corresponds to your chosen IBM Cloud region.

Available Regions​

IBM Cloud Object Storage is available in multiple regions worldwide:

Region CodeLocation
us-southDallas, USA
us-eastWashington DC, USA
eu-gbLondon, UK
eu-deFrankfurt, Germany
jp-tokTokyo, Japan
au-sydSydney, Australia

File Formats​

FormatDescription
jsonEach log entry is written as a separate JSON line (JSONL format)
multijsonAll log entries are written as a single JSON array
avroApache Avro format with schema
parquetApache Parquet columnar format with schema

Compression​

Some formats support built-in compression to reduce storage costs and transfer times. When supported, compression is applied at the file/block level before upload.

FormatDefaultCompression Codecs
JSON-Not supported
MultiJSON-Not supported
Avrozstddeflate, snappy, zstd
Parquetzstdgzip, snappy, zstd, brotli, lz4

File Management​

Files are rotated based on size (max_size parameter) or event count (batch_size parameter), whichever limit is reached first. Template variables in file names enable dynamic file naming for time-based partitioning.

Templates​

The following template variables can be used in file names:

VariableDescriptionExample
{{.Year}}Current year2024
{{.Month}}Current month01
{{.Day}}Current day15
{{.Timestamp}}Current timestamp in nanoseconds1703688533123456789
{{.Format}}File formatjson
{{.Extension}}File extensionjson
{{.Compression}}Compression typezstd
{{.TargetName}}Target namemy_logs
{{.TargetType}}Target typeibmcos
{{.Table}}Bucket namelogs

Multipart Upload​

Large files automatically use multipart upload protocol with configurable part size (part_size parameter). Default 5MB part size balances upload efficiency and memory usage.

Multiple Buckets​

Single target can write to multiple IBM COS buckets with different configurations, enabling data distribution strategies (e.g., raw data to one bucket, processed data to another).

Schema Requirements​

Avro and Parquet formats require schema definition files. Schema files must be accessible at the path specified in the schema parameter during target initialization.

Storage Classes​

IBM Cloud Object Storage supports multiple storage classes for cost optimization. Choose the appropriate class based on data access patterns and retention requirements.

Examples​

Basic Configuration​

The minimum configuration for a JSON IBM COS target:

targets:
- name: basic_ibm_cos
type: ibmcos
properties:
key: "0123456789abcdef0123456789abcdef"
secret: "fedcba9876543210fedcba9876543210fedcba98"
region: "us-south"
endpoint: "https://s3.us-south.cloud-object-storage.appdomain.cloud"
bucket: "datastream-logs"

Multiple Buckets​

Configuration for distributing data across multiple IBM COS buckets with different formats:

targets:
- name: multi_bucket_export
type: ibmcos
properties:
key: "0123456789abcdef0123456789abcdef"
secret: "fedcba9876543210fedcba9876543210fedcba98"
region: "eu-gb"
endpoint: "https://s3.eu-gb.cloud-object-storage.appdomain.cloud"
buckets:
- bucket: "raw-data-archive"
name: "raw-{{.Year}}-{{.Month}}-{{.Day}}.json"
format: "multijson"
compression: "gzip"
- bucket: "analytics-data"
name: "analytics-{{.Year}}/{{.Month}}/{{.Day}}/data_{{.Timestamp}}.parquet"
format: "parquet"
schema: "<schema definition>"
compression: "snappy"

Parquet Format​

Configuration for daily partitioned Parquet files:

targets:
- name: parquet_analytics
type: ibmcos
properties:
key: "0123456789abcdef0123456789abcdef"
secret: "fedcba9876543210fedcba9876543210fedcba98"
region: "jp-tok"
endpoint: "https://s3.jp-tok.cloud-object-storage.appdomain.cloud"
bucket: "analytics-lake"
name: "events/year={{.Year}}/month={{.Month}}/day={{.Day}}/part-{{.Timestamp}}.parquet"
format: "parquet"
schema: "<schema definition>"
compression: "snappy"
max_size: 536870912

High Reliability​

Configuration with enhanced settings:

targets:
- name: reliable_ibm_cos
type: ibmcos
pipelines:
- checkpoint
properties:
key: "0123456789abcdef0123456789abcdef"
secret: "fedcba9876543210fedcba9876543210fedcba98"
region: "us-east"
endpoint: "https://s3.us-east.cloud-object-storage.appdomain.cloud"
bucket: "critical-logs"
name: "logs-{{.Timestamp}}.json"
format: "json"
timeout: 60
part_size: 10

With Field Normalization​

Using field normalization for standard format:

targets:
- name: normalized_ibm_cos
type: ibmcos
properties:
key: "0123456789abcdef0123456789abcdef"
secret: "fedcba9876543210fedcba9876543210fedcba98"
region: "eu-de"
endpoint: "https://s3.eu-de.cloud-object-storage.appdomain.cloud"
bucket: "normalized-logs"
name: "logs-{{.Timestamp}}.json"
format: "json"
field_format: "cim"

Debug Configuration​

Configuration with debugging enabled:

targets:
- name: debug_ibm_cos
type: ibmcos
properties:
key: "0123456789abcdef0123456789abcdef"
secret: "fedcba9876543210fedcba9876543210fedcba98"
region: "au-syd"
endpoint: "https://s3.au-syd.cloud-object-storage.appdomain.cloud"
bucket: "test-logs"
name: "test-{{.Timestamp}}.json"
format: "json"
debug:
status: true
dont_send_logs: true