Skip to content

Full-line Index


What Is the Full-line Index

The Full-line Index is a log full-text search mode. Once enabled, the system includes all business fields in logs in the full-text index. When querying, you do not need to specify field names in advance; entering the value of any business field lets you search for logs containing that value.

For example, a log contains the following fields:

{
  "service": "order-api",
  "trace_id": "8d4f2a",
  "order_id": "O-1001",
  "amount": 128.5
}

After the Full-line Index is enabled, searching directly for 8d4f2a, O-1001, or 128.5 can find this log.

The Full-line Index applies to the following scenarios:

  • Logs have already been extracted into multiple structured fields and no longer rely on an entire message segment;
  • You need to quickly search across different business fields such as trace_id, order numbers, and user IDs;
  • The same log index contains multiple log structures, making it impossible to predetermine a unified search field for all logs;
  • You want to retain field-based query capabilities while also supporting full-text search across business fields.

Differences from the message-only Index

Comparison message-only Index Full-line Index
Full-text search scope message field only All business fields, including message if present
Whether original logs need to be retained message must be retained Not required; can be retained or deleted after field extraction
Suitable data Plain-text logs, unstructured logs JSON logs, structured logs extracted through Pipeline
Full-text index field message variant

The Full-line Index changes only the full-text search scope; it does not replace field filtering. For fields that have been extracted, you can still use field conditions such as service:order-api for precise filtering, aggregation, or analysis.

How It Works

The system uses variant to uniformly carry business fields in the Full-line Index and creates a full-text index only for variant.

When `message` does not exist:
business fields → `variant` → full-text index

When `message` exists:
`message` + other business fields → `variant` → full-text index

The specific rules are as follows:

  • The Full-line Index includes business fields in logs, but not system fields;
  • When a log does not have message, other business fields can still participate in full-text search;
  • When a log has message, the system does not delete this field but writes it to variant as an ordinary business field;
  • The system does not create two separate full-text indexes for message and variant at the same time;
  • The same log index can simultaneously store logs with and without message.

Whether message exists is determined by the original log content and the actual processing results of DataKit and Pipeline. Enabling the Full-line Index does not require deleting message.

Using or Switching to the Full-line Index

When you create a new log index, the Full-line Index is used by default and does not need to be enabled separately.

For existing log indexes that still use the message-only index field, perform the following steps:

  1. Go to Logs > Indexes;
  2. Edit the target log index;
  3. Expand Advanced Options;
  4. In Full-Text Index Field, select Full-line Index;
  5. Confirm the scope of impact and save the configuration.
Note

Switching from a message-only index to the Full-line Index is a one-way operation. After saving, you cannot switch back to the message-only index. Confirm related queries and data usage before performing this operation.

Prerequisites

The current workspace must support the Full-line Index. If this option is not shown on the page, check the workspace version and related feature permissions.

Querying with the Full-line Index

After the configuration is complete and logs have been written, go to Logs > Explorer and select the corresponding log index.

The Explorer supports text search, field filtering, combined search, JSON search, and DQL queries. For complete search syntax and usage instructions, see Explorer Search.

Enter the value of a business field in the search box to search across all business fields without specifying the field name.

Using the order log above as an example:

  • Entering O-1001 can match order_id;
  • Entering 8d4f2a can match trace_id;
  • Entering order-api can match service;
  • When message is retained in the log, you can also search for content in message.

Text search tokenizes the input. If you need to match complete, continuous content, wrap the search content in English double quotation marks. For more details, see Text Search.

Field Filtering

If you already know the field name, you can still use field conditions to narrow the query scope. For example:

service:order-api

Full-text search suits scenarios where you are not sure which field contains the content; field filtering suits scenarios where the field name is known and precise filtering or aggregation analysis is needed. The two can be used together. For complete field filtering syntax, see Filtering.

Query Notes

  • Full-line full-text search matches only business fields, not system fields;
  • When logs do not have message, the Explorer combines the current log's business fields to display log content;
  • System fields are not part of the content of Full-line Index logs;
  • If business fields have not yet been extracted from original logs, the Full-line Index can search only the fields that currently exist. Therefore, for JSON or plain-text logs, extract the required fields first as described below.

Extracting Log Fields

The Full-line Index indexes business fields that already exist in logs, but it does not automatically understand and split content inside message. If the original log is still an entire JSON object or plain text, you can use DataKit or Pipeline to extract the content into structured fields.

After field extraction, whether to retain the original message depends on actual needs:

  • If you need to view the complete original text or maintain compatibility with existing usage habits, you can retain message;
  • If there are enough structured fields and the original text no longer needs to be saved, you can delete message;
  • Regardless of whether message is retained, extracted business fields can participate in the Full-line full-text index.

Using DataKit to Extract JSON Fields

This applies to logs where each line is a standard JSON object and you want to directly extract all top-level fields. DataKit 2.9.0 and later can enable json_as_fields.

After enabling it, DataKit converts the top-level properties of the JSON root object into log fields after character decoding, ANSI cleanup, and multi-line joining, and then executes the Pipeline.

Host Log Collection

Edit logging.conf:

[[inputs.logging]]
  logfiles = ["/var/log/order/*.json"]
  source = "order"
  service = "order-api"
  json_as_fields = true

After saving the configuration, restart DataKit.

Kubernetes Container Log Collection

You can enable JSON field mode for a specified container through Pod Annotation:

metadata:
  annotations:
    datakit/order-api.logs: >-
      [{"source":"order","service":"order-api","json_as_fields":true}]

Here, order-api is the container name. You can also use the same JSON configuration in the container environment variable DATAKIT_LOGS_CONFIG.

Container Configuration Limits

json_as_fields currently supports only JSON log configuration in container environment variables and Pod Annotation/Label. It does not yet support the ClusterLoggingConfig CRD.

Log Streaming

Enable it in logstreaming.conf:

[inputs.logstreaming]
  json_as_fields = true

json_as_fields is a collector configuration, not an HTTP URL parameter. This configuration does not apply to the influxdb, firelens, and firehose types.

Extraction Results

Original log:

{"timestamp":"2026-08-18T10:00:00+08:00","level":"INFO","service":"order-api","trace_id":"8d4f2a","order_id":"O-1001","amount":128.5,"labels":{"channel":"web"}}

After successful conversion:

  • timestamp, level, service, trace_id, order_id, and amount become independent fields;
  • The labels object is saved as a compact JSON string;
  • DataKit no longer generates an additional message that stores the entire original JSON;
  • If the original JSON itself contains message, that field is retained normally.

Main field conversion rules:

  • Top-level strings, booleans, integers, and decimals preserve their types; objects and arrays are saved as compact JSON strings; null is ignored;
  • Invalid JSON, a root node that is not an object, or no valid fields falls back to the original message;
  • A . in field names is converted to _, newline characters are converted to spaces, and field names can be up to 256 bytes;
  • Each log retains at most 1024 fields, excluding tags;
  • JSON fields override same-name tags or fields from the collector;
  • time, source, date, and storage_index in JSON are renamed to json_time, json_source, json_date, and json_storage_index, respectively.

For complete rules, see JSON Field Mode.

Using Pipeline to Extract Fields

This applies to the following cases:

  • Only some fields need to be extracted;
  • Fields need to be renamed or converted;
  • Logs are not standard JSON;
  • You have an existing Pipeline cleaning flow that needs to be reused.

Extract All Top-Level Scalar Fields

DataKit 2.2.0 and later can use json_all():

# Raw data:
# {"service":"order-api","status":"info","trace_id":"8d4f2a","order_id":"O-1001","amount":128.5}

json_all(_, key_patterns=["*"])

After processing, you get independent fields such as service, status, trace_id, order_id, and amount.

json_all() extracts only top-level strings, numbers, and booleans. It does not recursively expand objects or arrays, nor does it save null. When neither include_keys nor key_patterns is configured, no fields are extracted. To extract all top-level scalar fields, you must explicitly set key_patterns=["*"].

Extract Only Specified Fields

json(_, service)
json(_, level, status)
json(_, trace_id)
json(_, order_id)
json(_, amount)

This approach lets you control the field scope that enters the Full-line Index and also unify fields from different logs under the same names.

Objects or arrays are not extracted by json_all(). If you need to retain such content, use json() with specified field extraction; the extraction result is saved as a JSON string.

Deleting message After Field Extraction

When extracting fields through Pipeline, the original log is still saved in message by default. If structured fields are sufficient for display and query needs, you can call drop_origin_data() after completing field extraction so that the original content is no longer saved.

json_all(_, key_patterns=["*"])
drop_origin_data()

You can also delete it after extracting specified fields:

json(_, service)
json(_, level, status)
json(_, trace_id)
json(_, order_id)
json(_, amount)
drop_origin_data()

If the JSON itself contains message and that field has also been extracted, you can keep or delete it as needed. If you are certain you do not need it, use:

json_all(_, key_patterns=["*"])
drop_origin_data()
drop_key(message)

The three functions serve different purposes:

  • drop_origin_data(): stops outputting the original text saved during initialization; the entire log is still uploaded;
  • drop_key(message): deletes the already extracted message field;
  • drop(): discards the entire log; the log is not uploaded and cannot be used to delete message.
Do Not Use drop() to Delete message

drop() marks the current log as to be discarded. After Pipeline execution finishes, the entire log is not uploaded.

For more syntax, see json(), json_all(), drop_origin_data(), and drop_key().

Verifying the Pipeline Locally

After saving the Pipeline script, you can run it on the host where DataKit resides:

datakit pipeline -P full_line_json.p \
  -T '{"service":"order-api","status":"info","trace_id":"8d4f2a","order_id":"O-1001","amount":128.5}'

Check whether the output meets the following expectations:

  • All fields that need to be searched have been extracted;
  • If drop_origin_data() is configured, the original message has been deleted;
  • The log is not marked as drop: true;
  • Fields such as status and log time meet business expectations.

FAQ

Does the Full-line Index require deleting message?

No. When message exists, the system writes it to variant as an ordinary business field and no longer creates a separate full-text index for message. Whether to delete message should be decided based on original content retention and display requirements.

After enabling the Full-line Index, why can I still not search JSON fields in message?

If the entire JSON is still only string content in message, the properties within it are not independent business fields. Although text in message can be searched, these properties cannot be directly used for field filtering or aggregation. It is recommended to use DataKit or Pipeline to extract the required fields first.

When there is no message, how does the log explorer display content?

When a log does not have message, the log explorer combines the current log's business fields to display log content; system fields are not part of the log content.

Can the same index contain both logs with and without message?

Yes. Both types of logs write business fields to variant according to unified rules, and variant participates in the full-text index.

Can the Full-line Index search system fields?

No. The Full-line Index contains only business fields, not system fields. System fields can still be queried through field filtering and other methods.

After switching to the Full-line Index, can I restore the message index?

No. Switching from a message-only index to the Full-line Index is a one-way operation and cannot be restored after saving. Existing message indexes that are not switched can continue to be used in the original way.

Feedback

Is this page helpful?