* [RFC PATCH v4 0/5] Add perf.data tracepoint events to trace.dat conversion
@ 2026-08-22 6:27 Tanushree Shah
2026-08-22 6:27 ` [RFC PATCH v4 1/5] perf/trace-dat: Add trace.dat export infrastructure Tanushree Shah
` (4 more replies)
0 siblings, 5 replies; 11+ messages in thread
From: Tanushree Shah @ 2026-08-22 6:27 UTC (permalink / raw)
To: acme, jolsa, adrian.hunter, vmolnaro, mpetlan, tmricht, maddy,
irogers, namhyung
Cc: linux-perf-users, linuxppc-dev, atrajeev, hbathini, Tejas.Manhas1,
Tanushree.Shah, rostedt, Tanushree Shah
This RFC patch series introduces support for converting perf.data files
containing tracepoint events into trace.dat format, enabling seamless
visualization and analysis using KernelShark.
======================
Background and Motivation
======================
Currently, perf and trace-cmd operate as separate tracing ecosystems with
incompatible data formats. Users who collect tracepoint data with
'perf record' cannot easily visualize it in KernelShark's graphical
timeline view or leverage trace-cmd's analysis capabilities.
This creates workflow friction when users need to:
- Visualize perf tracepoint data in KernelShark's interactive graphical
timeline
- Share trace data between perf and trace-cmd workflows and toolchains
- Perform architecture-independent conversion and analysis of traces
This conversion bridge eliminates these barriers by enabling seamless
data exchange between perf and trace-cmd ecosystems, allowing users to
choose the best tool for each analysis phase.
======================
Implementation Overview
======================
The series implements the trace.dat file format specification (version 7)
within perf's data conversion framework.
**Patch 1/5: Core trace.dat Export Infrastructure**
Introduces util/trace-dat.c and util/trace-dat.h implementing:
- Per-CPU raw event buffer management (init, collect, free)
- Ftrace ring buffer page construction
- trace.dat section writers (strings, options, flyrecord sections)
**Patch 2/5: Metadata Integration**
Extends util/trace-event-read.c to write trace.dat metadata during
perf.data
parsing:
- Initial format header (magic, version, endian, page size, compression)
- Section 16: HEADER INFO (header_page + header_event)
- Section 17: FTRACE EVENT FORMATS
- Section 18: EVENT FORMATS (per system/event format files)
- Section 19: KALLSYMS
- Section 21: CMDLINES
- Section 15: STRINGS (written last after all sections)
**Patch 3/5: Conversion Backend**
Implements util/data-convert-trace.c with trace_convert__perf2dat()
function:
- Processes PERF_TYPE_TRACEPOINT samples via process_sample_event()
- Collects raw event data per-CPU using trace_dat__collect_cpu_event()
- Writes OPTIONS sections (CPUCOUNT, TRACECLOCK, metadata offsets)
- Writes FLYRECORD section with per-CPU ring buffer pages
**Patch 4/5: User Interface**
Extends tools/perf/builtin-data.c with --to-trace-dat option:
- Adds command-line option for trace.dat output
- Mutually exclusive with --to-ctf and --to-json
- Calls trace_convert__perf2dat() to perform conversion
**Patch 5/5: Shell Test**
Adds a shell test (tools/perf/tests/shell/) covering normal tracepoint
recordings, pipe mode and mixed tracepoint/non-tracepoint recording
conversions and --force flag behaviour.
======================
Current Implementation Details
======================
**trace.dat Format Version:**
The implementation currently targets trace.dat format version 7, which
is the stable version supported by current trace-cmd releases (v3.x).
This version is hardcoded to ensure compatibility with existing
trace-cmd and KernelShark installations. Future enhancements could add
version negotiation or support for newer format versions as they become
standardized.
**Compression Strategy:**
Compression is explicitly disabled (set to NONE) in the generated
trace.dat files.
This design choice:
- Simplifies the initial implementation and testing
- Ensures maximum compatibility across trace-cmd versions
- Avoids external compression library dependencies
Future work could add support for various compression algorithms (zlib,
zstd, lz4) with runtime selection via command-line options, significantly
reducing file sizes for large traces.
======================
Usage Example
======================
```bash
*Record tracepoint events with perf*
perf record -e sched:sched_switch -e sched:sched_wakeup -a sleep 10
*Convert to trace.dat format*
perf data convert --to-trace-dat=output.dat
*Verify trace.dat structure*
trace-cmd dump --summary output.dat
*Analyze with trace-cmd*
trace-cmd report output.dat
*Visualize in KernelShark*
kernelshark output.dat
```
**Conversion Output:**
```
[ perf data convert: Converted 'perf.data' into trace.dat format
'output.dat' ]
[ perf data convert: Converted 2684 events ]
```
**trace-cmd dump --summary Output:**
```
Tracing meta data in file output.dat:
[Initial format]
7 [Version]
0 [Little endian]
8 [Bytes in a long]
65536 [Page size, bytes]
none [Compression algorithm]
[Compression version]
[buffer "", "local" clock, 65536 page size, 16 cpus, 1048576 bytes
flyrecord data]
[10 options]
[Saved command lines, 0 bytes]
[Kallsyms, 0 bytes]
[Ftrace format, 0 events]
[Header page, 206 bytes]
[Header event, 205 bytes]
[Events format, 1 systems]
[9 sections]
```
======================
Testing and Verification
======================
The series has been extensively tested with:
- Various tracepoint events (sched, irq, syscalls, block I/O)
- Mixed recordings containing both tracepoint and non-tracepoint events
(only tracepoints converted)
- Verification with trace-cmd report and KernelShark visualization
- Memory leak testing with Valgrind (0 bytes leaked)
- Cross-architecture testing: v1 tested x86_64 and ppc64le. v2 adds
s390 (big-endian) perf.data converted on both ppc64le
(little-endian) and x86_64 (little-endian) hosts, in addition to
same-arch x86_64 (LE->LE) and ppc64le (BE->BE) conversion.
- Pipe mode support has been tested end-to-end. (in v2)
- Time filtering: verified --time option correctly limits converted
events to the requested range, consistent with --to-json behaviour.
(in v4)
All generated trace.dat files successfully open in:
- trace-cmd report (v3.1+)
- KernelShark (v2.0+)
======================
Next Steps
======================
We would highly appreciate reviews, comments, and feedback on:
- The overall architectural approach and integration points
- Compatibility considerations with trace-cmd ecosystem
- Performance characteristics for large-scale traces
- Additional use cases or workflow scenarios
- Future enhancement priorities
---
Changes in v4
Addressing Sashiko AI review findings on v3:
Timestamp correctness (trace-dat.c):
- Fix TIME_EXTEND delta_upper shift: >> 5 (TRACE_DAT_RECORD_TIME_SHIFT)
should be >> 27, causing wildly inflated timestamps in trace-cmd.
- Advance base_ts after every event so time_delta is relative to the
preceding record, not a stale page base.
- Introduce page_base_ts updated only at page boundaries, keeping the
page header timestamp correct independently of per-event base_ts.
Memory safety (trace-dat.c):
- Fix leak on TIME_EXTEND calloc() failure: use goto out_free instead
of return -ENOMEM so page_records/page_rec_sizes are freed.
I/O correctness (trace-event-read.c):
- Check fseek() return values when patching the event formats section
size; set trace_dat_write_failed on any failure to prevent silent
corruption of subsequent writes.
Resource management (data-convert-trace.c):
- Fix orphaned 0-byte output file when fdopen() fails in !opts->force
path: use goto out_close so unlink() is called on error.
- Honor --time filtering using perf_time__parse_for_ranges() and
perf_time__ranges_skip_sample(), consistent with JSON/CTF converters.
- Add explicit <stdio.h> include for fdopen()/fopen()/fclose().
Header (trace-dat.h):
- Add explicit <stdint.h> for uint16_t/uint32_t/uint64_t on musl libc.
Shell test (patch 5):
- Replace 'cycles' with 'cpu-clock -a' for portability on s390, KVM
guests and unprivileged containers without hardware PMU support.
- Add -a to all perf record/sleep invocations per Ian's suggestion.
- Fix EXIT trap to pass through exit code 2 (skip) unchanged.
- Move second mktemp after trap registration to avoid temp file leak.
- Drop rm -f "$result" before converter in test_trace_converter_command;
--force handles overwriting and deleting first opens a symlink race.
- Move rm -f "$result" to before perf record in pipe and mixed-events
tests to prevent stale output being validated on early return.
Changes in v3
- Rebase on latest perf-tools-next and resolve merge conflicts.
- Drop evsel parameter from process_sample_event() following upstream
commit "perf tool: Remove evsel from tool APIs that pass the sample".
Changes in v2
Addressing the Sashiko AI review findings on v1:
Cross-arch correctness:
- Introduce to_file_u16/u32/u64 helpers (wrapping tep_read_number())
to write all multi-byte fields in the recorded machine's byte order;
apply throughout metadata sections and flyrecord page/record headers
(ts, commit, TIME_EXTEND, large-event data_len).
- Fix flyrecord record header bit layout for big-endian files.
The record header word bit layout differs by file endianness,
matching kbuffer-parse.c type_len4host()/ts4host():
LE: type_len in bits [4:0], time_delta in bits [31:5]
BE: type_len in bits [31:27], time_delta in bits [26:0]
Pipe mode:
- Add process_attr(), process_feature(), and process_tracing_data()
callbacks required for pipe mode operation.
- Defer CPU buffer initialisation until the first tracepoint sample,
after process_feature()/process_tracing_data() have populated the
session header. This ensures the recorded machine's CPU count is
used rather than the host's - critical for cross-platform analysis.
Format compliance:
- Implement TIME_EXTEND records for timestamp deltas >27 bits to
prevent silent truncation and maintain chronological ordering.
- Fix large event encoding (>=29 words): use type_len=0 with a
separate 32-bit length word, avoiding collision with reserved types
(PADDING=29, TIME_EXTEND=30, TIME_STAMP=31).
- Add bounds check rejecting records larger than a page payload before
batching, preventing heap overflow in trace_dat__write_page().
- Fix flyrecord section_size to exclude the 16-byte section header,
matching trace.dat specification and trace-cmd behaviour.
CLI behavior:
- Fix --force flag: open with O_CREAT|O_EXCL when force is not set,
failing with -EEXIST instead of silently overwriting existing files.
Memory safety:
- Fix realloc overwrite of cpu_events->events and page_records on
failure: use temporary pointers, only commit on success.
- Fix use-after-free/double-free in sequential page_records realloc
failure: replace with malloc+memcpy+free pattern.
- Fix section_size computed from before section header position.
- Add NULL checks for get_tracing_file(), calloc() padding, and
trace_dat_options_offset assignment on write failure.
- Use goto out_free on record allocation failure to avoid leaking
accumulated page_records entries.
- Replace direct read() with do_read() in read_proc_kallsyms() to
handle short reads correctly.
- On fwrite failure, set trace_dat_write_failed and continue parsing
so that perf.data processing completes normally.
Testing (new in v2):
- Add shell test covering conversion, trace-cmd dump validation,
sched_switch event verification, and --force flag behaviour.
Documentation (new in v2):
- Add documentation for 'perf data convert --to-trace-dat', covering
usage and supported options.
v1: https://lore.kernel.org/linux-perf-users/20260608125951.90425-2-tshah@linux.ibm.com/
Tanushree Shah (5):
perf/trace-dat: Add trace.dat export infrastructure
perf/trace-event: Write trace.dat metadata sections during parsing
perf data-convert: Add perf.data to trace.dat conversion backend
perf data: Add --to-trace-dat option for converting perf.data
tracepoint events into trace.dat format
perf test: Add test validating trace.dat generated by 'perf data
convert --to-trace-dat'
tools/perf/Documentation/perf-data.txt | 7 +
tools/perf/builtin-data.c | 43 +
...rf_data_converter_tracepoints_trace_dat.sh | 173 ++++
tools/perf/util/Build | 2 +
tools/perf/util/data-convert-trace.c | 264 ++++++
tools/perf/util/data-convert.h | 4 +
tools/perf/util/trace-dat.c | 883 ++++++++++++++++++
tools/perf/util/trace-dat.h | 114 +++
tools/perf/util/trace-event-read.c | 307 +++++-
9 files changed, 1790 insertions(+), 7 deletions(-)
create mode 100755 tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh
create mode 100644 tools/perf/util/data-convert-trace.c
create mode 100644 tools/perf/util/trace-dat.c
create mode 100644 tools/perf/util/trace-dat.h
--
2.47.1
^ permalink raw reply [flat|nested] 11+ messages in thread
* [RFC PATCH v4 1/5] perf/trace-dat: Add trace.dat export infrastructure
2026-08-22 6:27 [RFC PATCH v4 0/5] Add perf.data tracepoint events to trace.dat conversion Tanushree Shah
@ 2026-08-22 6:27 ` Tanushree Shah
2026-08-22 6:40 ` sashiko-bot
2026-08-22 6:27 ` [RFC PATCH v4 2/5] perf/trace-event: Write trace.dat metadata sections during parsing Tanushree Shah
` (3 subsequent siblings)
4 siblings, 1 reply; 11+ messages in thread
From: Tanushree Shah @ 2026-08-22 6:27 UTC (permalink / raw)
To: acme, jolsa, adrian.hunter, vmolnaro, mpetlan, tmricht, maddy,
irogers, namhyung
Cc: linux-perf-users, linuxppc-dev, atrajeev, hbathini, Tejas.Manhas1,
Tanushree.Shah, rostedt, Tanushree Shah
Add new utility files util/trace-dat.c and util/trace-dat.h
implementing the infrastructure for exporting perf.data tracepoints
to trace.dat format compatible with trace-cmd and KernelShark.
trace-dat.c defines all globals and functions needed for:
- Per-cpu raw event buffer management (init_cpu_buffers,
collect_cpu_event, free_cpu_buffers).
- ftrace ring buffer page construction (write_page, write_cpu_dat),
including TIME_EXTEND records for large timestamp deltas.
- trace.dat section writers (write_strings_section,
write_options_section1, write_options_section2,
write_flyrecord_section).
trace-dat.h declares all globals and function prototypes to be
used by data-convert-trace.c and trace-event-read.c.
Signed-off-by: Tanushree Shah <tshah@linux.ibm.com>
---
tools/perf/util/Build | 1 +
tools/perf/util/trace-dat.c | 825 ++++++++++++++++++++++++++++++++++++
tools/perf/util/trace-dat.h | 83 ++++
3 files changed, 909 insertions(+)
create mode 100644 tools/perf/util/trace-dat.c
create mode 100644 tools/perf/util/trace-dat.h
diff --git a/tools/perf/util/Build b/tools/perf/util/Build
index b26a0b1ddfa3..12f5bbe32770 100644
--- a/tools/perf/util/Build
+++ b/tools/perf/util/Build
@@ -100,6 +100,7 @@ perf-util-y += trace-event-scripting.o
perf-util-$(CONFIG_LIBTRACEEVENT) += trace-event.o
perf-util-$(CONFIG_LIBTRACEEVENT) += trace-event-parse.o
perf-util-$(CONFIG_LIBTRACEEVENT) += trace-event-read.o
+perf-util-$(CONFIG_LIBTRACEEVENT) += trace-dat.o
perf-util-y += sort.o
perf-util-y += hist.o
perf-util-y += util.o
diff --git a/tools/perf/util/trace-dat.c b/tools/perf/util/trace-dat.c
new file mode 100644
index 000000000000..55a7bd982c6f
--- /dev/null
+++ b/tools/perf/util/trace-dat.c
@@ -0,0 +1,825 @@
+// SPDX-License-Identifier: GPL-2.0-or-later
+/*
+ * Copyright 2026, IBM Corporation
+ * Author: Tanushree Shah <tshah@linux.ibm.com>
+ *
+ * trace-dat.c
+ *
+ * This file implements the trace.dat format writer for perf tool.
+ * It collects trace events from multiple CPUs and writes them in
+ * the trace-cmd compatible format.
+ */
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <unistd.h>
+#include <errno.h>
+#include "api/fs/tracing_path.h"
+#include "trace-dat.h"
+#include "trace-event.h"
+#include "session.h"
+#include "header.h"
+#include "../perf.h"
+#include "debug.h"
+
+/* ftrace ring buffer constants for trace.dat flyrecord section
+ *
+ * Each page has a 16-byte header (timestamp + commit size), followed by
+ * variable-length records. Each record has a 4-byte header word encoding:
+ * Bits 0-4: Type/Length field (5 bits, masked by TYPE_LEN_MASK)
+ * Bits 5-31: Time delta from page base timestamp (27 bits, masked by TIME_MASK)
+ */
+#define TRACE_DAT_RECORD_HEADER_SIZE 16 /* Page header: 8-byte ts + 8-byte commit */
+#define TRACE_DAT_RECORD_TYPE_LEN_MASK 0x1F /* Extract lower 5 bits for type/length */
+#define TRACE_DAT_RECORD_TIME_SHIFT 5 /* Shift to extract time delta */
+#define TRACE_DAT_RECORD_TIME_MASK 0x07FFFFFF /* Mask for 27-bit time delta */
+#define TRACE_DAT_WORD_SIZE 4 /* Records aligned to 4-byte boundaries */
+#define TRACE_DAT_WORD_ALIGN_MASK 3
+
+/* Initial capacity for per-CPU event buffer (grows by doubling) */
+#define INITIAL_EVENT_CAPACITY 1024
+/* Initial capacity for page record array (grows by doubling) */
+#define INITIAL_PAGE_RECORD_CAPACITY 64
+/* Buffer size for reading trace_clock string from debugfs/tracefs */
+#define CLOCK_BUFFER_SIZE 256
+
+FILE *trace_dat_fp;
+int trace_dat_page_size;
+int trace_dat_nr_cpus;
+long trace_dat_options_offset;
+long trace_dat_header_info_offset;
+long trace_dat_events_format_offset;
+long trace_dat_ftrace_format_offset;
+long trace_dat_kallsyms_offset;
+long trace_dat_cmdline_offset;
+long trace_dat_next_options_offset;
+
+
+/**
+ * struct cpu_event - Single trace event from a CPU
+ * @ts: Timestamp of the event
+ * @raw: Raw event data
+ * @raw_size: Size of raw event data in bytes
+ */
+struct cpu_event {
+ unsigned long long ts;
+ void *raw;
+ unsigned int raw_size;
+};
+
+/**
+ * struct cpu_events - Collection of trace events for a single CPU
+ * @events: Array of events
+ * @count: Number of events currently stored
+ * @capacity: Maximum number of events that can be stored
+ */
+struct cpu_events {
+ struct cpu_event *events;
+ int count;
+ int capacity;
+};
+
+static struct cpu_events *trace_cpu_data;
+static long *buffer_opt_cpu_offsets_pos;
+static long opt_payload_start;
+
+/* Allocate per-cpu event buffers for tracepoint data collection */
+int trace_dat__init_cpu_buffers(int nr_cpus)
+{
+ trace_cpu_data = calloc(nr_cpus, sizeof(struct cpu_events));
+ if (!trace_cpu_data)
+ return -ENOMEM;
+ buffer_opt_cpu_offsets_pos = calloc(nr_cpus, sizeof(long));
+ if (!buffer_opt_cpu_offsets_pos) {
+ free(trace_cpu_data);
+ trace_cpu_data = NULL;
+ return -ENOMEM;
+ }
+ trace_dat_nr_cpus = nr_cpus;
+ return 0;
+}
+
+/* Store raw tracepoint event data in per-cpu buffer for trace.dat
+ * flyrecord
+ */
+int trace_dat__collect_cpu_event(int cpu, unsigned long long ts,
+ void *raw, unsigned int raw_size)
+{
+ struct cpu_events *cpu_events;
+ void *raw_copy;
+
+ if (!trace_cpu_data || cpu < 0 || cpu >= trace_dat_nr_cpus)
+ return -EINVAL;
+
+ if (!raw || raw_size == 0)
+ return -EINVAL;
+
+ cpu_events = &trace_cpu_data[cpu];
+
+ if (cpu_events->count >= cpu_events->capacity) {
+ int new_capacity = cpu_events->capacity ?
+ cpu_events->capacity * 2 : INITIAL_EVENT_CAPACITY;
+ struct cpu_event *tmp;
+
+ tmp = realloc(cpu_events->events, new_capacity * sizeof(*cpu_events->events));
+ if (!tmp)
+ return -ENOMEM;
+ cpu_events->events = tmp;
+ cpu_events->capacity = new_capacity;
+ }
+
+ raw_copy = malloc(raw_size);
+ if (!raw_copy)
+ return -ENOMEM;
+
+ memcpy(raw_copy, raw, raw_size);
+ cpu_events->events[cpu_events->count].ts = ts;
+ cpu_events->events[cpu_events->count].raw = raw_copy;
+
+ cpu_events->events[cpu_events->count].raw_size = raw_size;
+ cpu_events->count++;
+
+ return 0;
+}
+
+/* Write a single page of trace records */
+static int trace_dat__write_page(FILE *fp, unsigned long long base_ts,
+ char **records, int *rec_sizes, int nr_recs)
+{
+ unsigned long long commit = 0;
+ int offset = TRACE_DAT_RECORD_HEADER_SIZE;
+ int i;
+ char *page;
+
+ page = calloc(1, trace_dat_page_size);
+ if (!page)
+ return -ENOMEM;
+
+ for (i = 0; i < nr_recs; i++) {
+ memcpy(page + offset, records[i], rec_sizes[i]);
+ offset += rec_sizes[i];
+ commit += rec_sizes[i];
+ }
+
+ memcpy(page, &base_ts, sizeof(base_ts));
+ memcpy(page + sizeof(base_ts), &commit, sizeof(commit));
+
+ if (!fwrite(page, 1, trace_dat_page_size, fp)) {
+ free(page);
+ return -EIO;
+ }
+ free(page);
+
+ return 0;
+}
+
+/* Write all trace data for a single CPU as trace.dat flyrecord pages */
+static int trace_dat__write_cpu_dat(FILE *fp, int cpu, unsigned long long *file_offset_out)
+{
+ struct cpu_events *cpu_events = &trace_cpu_data[cpu];
+ unsigned long long base_ts;
+ unsigned long long file_offset;
+ char **page_records = NULL;
+ int *page_rec_sizes = NULL;
+ int page_cap = 0;
+ int nr_page_recs = 0;
+ int page_size_used = 0;
+ int ret = 0;
+ int i, j;
+
+ file_offset = ftell(fp);
+ *file_offset_out = file_offset;
+
+ if (cpu_events->count == 0) {
+ char *empty_page = calloc(1, trace_dat_page_size);
+
+ if (!empty_page)
+ return -ENOMEM;
+ if (!fwrite(empty_page, 1, trace_dat_page_size, fp)) {
+ free(empty_page);
+ return -EIO;
+ }
+ free(empty_page);
+ return 0;
+ }
+
+ base_ts = cpu_events->events[0].ts;
+
+ for (i = 0; i < cpu_events->count; i++) {
+ struct cpu_event *event = &cpu_events->events[i];
+ unsigned long long time_delta = event->ts - base_ts;
+ unsigned int data_len = event->raw_size;
+ unsigned int words;
+ unsigned int type_len;
+ unsigned int payload_offset;
+ unsigned int hdr_word;
+ char *extend = NULL;
+ int extend_size = 0;
+ char *data_rec;
+ int data_rec_size;
+ int needed_size;
+
+ /* Emit TIME_EXTEND when delta does not fit in 27 bits */
+ if (time_delta > TRACE_DAT_RECORD_TIME_MASK) {
+ unsigned int extend_hdr;
+ unsigned int delta_upper;
+
+ extend_size = TRACE_DAT_RECORD_TIME_EXTEND_SIZE;
+ extend = calloc(1, extend_size);
+ if (!extend)
+ return -ENOMEM;
+
+ extend_hdr =
+ ((time_delta & TRACE_DAT_RECORD_TIME_MASK) <<
+ TRACE_DAT_RECORD_TIME_SHIFT) |
+ TRACE_DAT_RECORD_TYPE_TIME_EXTEND;
+ delta_upper = time_delta >> TRACE_DAT_RECORD_TIME_SHIFT;
+
+ memcpy(extend, &extend_hdr, TRACE_DAT_WORD_SIZE);
+ memcpy(extend + TRACE_DAT_WORD_SIZE, &delta_upper,
+ TRACE_DAT_WORD_SIZE);
+
+ time_delta = 0;
+ }
+ words = (data_len + TRACE_DAT_WORD_ALIGN_MASK) / TRACE_DAT_WORD_SIZE;
+ /*
+ * Encode event length per ftrace ring buffer format:
+ * - Small events (≤28 words): type_len = word count
+ * - Large events (≥29 words): type_len = 0, followed by byte length
+ */
+ if (words <= TRACE_DAT_RECORD_TYPE_LEN_MAX) {
+ type_len = words;
+ payload_offset = TRACE_DAT_WORD_SIZE;
+ data_rec_size = TRACE_DAT_WORD_SIZE + data_len;
+ } else {
+ type_len = 0;
+ payload_offset = 2 * TRACE_DAT_WORD_SIZE;
+ data_rec_size = 2 * TRACE_DAT_WORD_SIZE + data_len;
+ }
+
+ if (data_rec_size % TRACE_DAT_WORD_SIZE)
+ data_rec_size += TRACE_DAT_WORD_SIZE -
+ (data_rec_size % TRACE_DAT_WORD_SIZE);
+
+ /* Reject records that cannot fit in a single page payload */
+ if (data_rec_size > trace_dat_page_size - TRACE_DAT_RECORD_HEADER_SIZE) {
+ free(extend);
+ ret = -E2BIG;
+ goto out_free;
+ }
+
+ needed_size = extend_size + data_rec_size;
+
+ /* Check page fit BEFORE allocating data record */
+ if (page_size_used + needed_size >
+ trace_dat_page_size - TRACE_DAT_RECORD_HEADER_SIZE) {
+ ret = trace_dat__write_page(fp, base_ts,
+ page_records, page_rec_sizes,
+ nr_page_recs);
+
+ for (j = 0; j < nr_page_recs; j++)
+ free(page_records[j]);
+
+ nr_page_recs = 0;
+ page_size_used = 0;
+ base_ts = event->ts;
+
+ if (ret < 0) {
+ free(extend);
+ goto out_free;
+ }
+
+ /*
+ * New page base matches this event timestamp, so no
+ * TIME_EXTEND is needed anymore.
+ */
+ free(extend);
+ extend = NULL;
+ extend_size = 0;
+ time_delta = 0;
+ }
+
+ hdr_word = (time_delta << TRACE_DAT_RECORD_TIME_SHIFT) | type_len;
+
+ data_rec = calloc(1, data_rec_size);
+ if (!data_rec) {
+ free(extend);
+ ret = -ENOMEM;
+ goto out_free;
+ }
+
+ memcpy(data_rec, &hdr_word, TRACE_DAT_WORD_SIZE);
+
+ /* Large events: write actual byte length after header */
+ if (type_len == 0)
+ memcpy(data_rec + TRACE_DAT_WORD_SIZE, &data_len, TRACE_DAT_WORD_SIZE);
+
+ memcpy(data_rec + payload_offset, event->raw, data_len);
+
+ if (nr_page_recs + (extend ? 1 : 0) >= page_cap) {
+ char **new_records;
+ int *new_sizes;
+ int needed = nr_page_recs + (extend ? 1 : 0) + 1;
+ int new_capacity = page_cap ? page_cap * 2 :
+ INITIAL_PAGE_RECORD_CAPACITY;
+
+ while (new_capacity < needed)
+ new_capacity *= 2;
+
+ new_records = malloc(new_capacity * sizeof(*new_records));
+ new_sizes = malloc(new_capacity * sizeof(*new_sizes));
+ if (!new_records || !new_sizes) {
+ free(new_records);
+ free(new_sizes);
+ free(extend);
+ free(data_rec);
+ ret = -ENOMEM;
+ goto out_free;
+ }
+
+ if (nr_page_recs > 0) {
+ memcpy(new_records, page_records,
+ nr_page_recs * sizeof(*new_records));
+ memcpy(new_sizes, page_rec_sizes,
+ nr_page_recs * sizeof(*new_sizes));
+ }
+
+ free(page_records);
+ free(page_rec_sizes);
+
+ page_records = new_records;
+ page_rec_sizes = new_sizes;
+ page_cap = new_capacity;
+ }
+
+ if (extend) {
+ page_records[nr_page_recs] = extend;
+ page_rec_sizes[nr_page_recs] = extend_size;
+ nr_page_recs++;
+ page_size_used += extend_size;
+ }
+
+ page_records[nr_page_recs] = data_rec;
+ page_rec_sizes[nr_page_recs] = data_rec_size;
+ nr_page_recs++;
+ page_size_used += data_rec_size;
+ }
+
+ if (nr_page_recs > 0) {
+ ret = trace_dat__write_page(fp, base_ts,
+ page_records, page_rec_sizes, nr_page_recs);
+ }
+out_free:
+ for (j = 0; j < nr_page_recs; j++)
+ free(page_records[j]);
+ free(page_records);
+ free(page_rec_sizes);
+ return ret;
+}
+
+/* Write the strings section containing section name lookup table */
+int trace_dat__write_strings_section(void)
+{
+ unsigned short section_id = TRACE_DAT_SECTION_STRINGS;
+ unsigned short flags = 0;
+ unsigned long long section_size = 0;
+ static const char * const section_names[] = {
+ "headers", /* offset 0 - strid for section 16 */
+ "ftrace event formats", /* offset 8 - strid for section 17 */
+ "events format", /* offset 29 - strid for section 18 */
+ "kallsyms", /* offset 43 - strid for section 19 */
+ "cmdlines", /* offset 52 - strid for section 21 */
+ "strings", /* offset 61 - strid for section 15 */
+ "options", /* offset 69 - strid for options 1 */
+ "options", /* offset 77 - strid for options 2 */
+ "buffer-flyrecord", /* offset 85 - strid for flyrecord 3 */
+ NULL
+ };
+
+ /* string_id points to "strings" string itself */
+ unsigned int string_id = STRID_STRINGS;
+ int i;
+
+ if (!trace_dat_fp)
+ return -EBADF;
+
+ for (i = 0; section_names[i] != NULL; i++)
+ section_size += strlen(section_names[i]) + 1;
+
+ /* write section header */
+ if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp) ||
+ !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp))
+ return -EIO;
+
+ /* write strings */
+ for (i = 0; section_names[i] != NULL; i++)
+ if (!fwrite(section_names[i], 1, strlen(section_names[i]) + 1, trace_dat_fp))
+ return -EIO;
+ return 0;
+}
+
+/* Writes options section containing CPUCOUNT, TRACECLOCK, EVENT_FORMAT, HEADER_INFO,
+ * FTRACE_EVENTS, KALLSYMS, CMDLINES options, ending with DONE option pointing to next section.
+ */
+int trace_dat__write_options_section1(void)
+{
+ unsigned short section_id = TRACE_DAT_SECTION_OPTIONS;
+ unsigned short flags = 0;
+ unsigned int string_id = STRID_OPTIONS_1;
+ unsigned long long section_size = 0;
+ long section_size_pos;
+ long payload_start;
+ unsigned long long section_start;
+ unsigned short opt_id;
+ unsigned int opt_size;
+ char clock_buf[CLOCK_BUFFER_SIZE];
+ FILE *clock_file;
+ size_t bytes_read;
+ char *path;
+ unsigned long long next_offset;
+ long end_pos;
+
+ if (!trace_dat_fp)
+ return -EBADF;
+
+ /* fill options_offset in initial format */
+ section_start = ftell(trace_dat_fp);
+
+ if (fseek(trace_dat_fp, trace_dat_options_offset, SEEK_SET) < 0 ||
+ !fwrite(§ion_start, sizeof(unsigned long long), 1, trace_dat_fp) ||
+ fseek(trace_dat_fp, 0, SEEK_END) < 0)
+ return -EIO;
+
+ /* write section header */
+ if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp))
+ return -EIO;
+ section_size_pos = ftell(trace_dat_fp);
+ if (!fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp))
+ return -EIO;
+
+ payload_start = ftell(trace_dat_fp);
+
+ /* CPUCOUNT option */
+ opt_id = TRACE_DAT_OPTION_CPUCOUNT;
+ opt_size = sizeof(unsigned int);
+
+ if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) ||
+ !fwrite(&trace_dat_nr_cpus, sizeof(unsigned int), 1, trace_dat_fp))
+ return -EIO;
+
+ /* TRACECLOCK option */
+ opt_id = TRACE_DAT_OPTION_TRACECLOCK;
+
+ path = get_tracing_file("trace_clock");
+ if (path) {
+ clock_file = fopen(path, "r");
+ put_tracing_file(path);
+ } else {
+ clock_file = NULL;
+ }
+ if (clock_file) {
+ bytes_read = fread(clock_buf, 1, sizeof(clock_buf) - 1, clock_file);
+ fclose(clock_file);
+ clock_buf[bytes_read] = '\0';
+ } else {
+ strcpy(clock_buf, "local\n");
+ bytes_read = strlen(clock_buf);
+ }
+ opt_size = bytes_read + 1;
+ if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) ||
+ !fwrite(clock_buf, 1, opt_size, trace_dat_fp))
+ return -EIO;
+
+ /* EVENT option */
+ opt_id = TRACE_DAT_OPTION_EVENT;
+ opt_size = sizeof(unsigned long long);
+
+ if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) ||
+ !fwrite(&trace_dat_events_format_offset, sizeof(unsigned long long),
+ 1, trace_dat_fp))
+ return -EIO;
+
+ /* HEADER option */
+ opt_id = TRACE_DAT_OPTION_HEADER;
+ opt_size = sizeof(unsigned long long);
+
+ if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) ||
+ !fwrite(&trace_dat_header_info_offset, sizeof(unsigned long long),
+ 1, trace_dat_fp))
+ return -EIO;
+
+ /* FTRACE option */
+ opt_id = TRACE_DAT_OPTION_FTRACE;
+ opt_size = sizeof(unsigned long long);
+
+ if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) ||
+ !fwrite(&trace_dat_ftrace_format_offset, sizeof(unsigned long long),
+ 1, trace_dat_fp))
+ return -EIO;
+
+ /* KALLSYMS option */
+ opt_id = TRACE_DAT_OPTION_KALLSYMS;
+ opt_size = sizeof(unsigned long long);
+
+ if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) ||
+ !fwrite(&trace_dat_kallsyms_offset, sizeof(unsigned long long),
+ 1, trace_dat_fp))
+ return -EIO;
+
+ /* CMDLINE option */
+ opt_id = TRACE_DAT_OPTION_CMDLINE;
+ opt_size = sizeof(unsigned long long);
+
+ if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) ||
+ !fwrite(&trace_dat_cmdline_offset, sizeof(unsigned long long),
+ 1, trace_dat_fp))
+ return -EIO;
+
+ /* DONE option id - next_options_offset filled later */
+ opt_id = TRACE_DAT_OPTION_DONE;
+ opt_size = sizeof(unsigned long long);
+ next_offset = 0; /* placeholder */
+
+ trace_dat_next_options_offset = ftell(trace_dat_fp);
+ if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) ||
+ !fwrite(&next_offset, sizeof(unsigned long long), 1, trace_dat_fp))
+ return -EIO;
+
+ /* fill section size */
+ end_pos = ftell(trace_dat_fp);
+
+ section_size = end_pos - payload_start;
+ if (fseek(trace_dat_fp, section_size_pos, SEEK_SET) < 0 ||
+ !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp) ||
+ fseek(trace_dat_fp, end_pos, SEEK_SET) < 0)
+ return -EIO;
+
+ return 0;
+
+}
+
+/* Writes options section containing BUFFER option with flyrecord section
+ * (flyrecord section offset, clock type, page size, CPU count,
+ * per-CPU offsets/sizes) and DONE option.
+ */
+int trace_dat__write_options_section2(void)
+{
+ unsigned short section_id = TRACE_DAT_SECTION_OPTIONS;
+ unsigned short flags = 0;
+ unsigned int string_id = STRID_OPTIONS_2;
+ unsigned long long section_size = 0;
+ long section_size_pos;
+ long payload_start;
+ int cpu;
+ unsigned short opt_id = TRACE_DAT_OPTION_BUFFER;
+ unsigned int opt_size = 0;
+ long opt_size_pos;
+ unsigned long long data_offset = 0;
+ unsigned int page_size = (unsigned int)trace_dat_page_size;
+ const char *clock = "local";
+ unsigned long long next;
+ long end_pos;
+ unsigned long long cpu_offset;
+ unsigned long long cpu_size;
+ unsigned short done_id;
+ unsigned int done_size;
+
+ if (!trace_dat_fp)
+ return -EINVAL;
+
+ /* fill done1 next offset - points to this section */
+ next = ftell(trace_dat_fp);
+
+ if (fseek(trace_dat_fp, trace_dat_next_options_offset + 2 + 4, SEEK_SET) < 0 ||
+ !fwrite(&next, sizeof(unsigned long long), 1, trace_dat_fp) ||
+ fseek(trace_dat_fp, 0, SEEK_END) < 0)
+ return -EIO;
+
+ /* write section header */
+ if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp))
+ return -EIO;
+ section_size_pos = ftell(trace_dat_fp);
+ if (!fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp))
+ return -EIO;
+
+ payload_start = ftell(trace_dat_fp);
+
+ /* BUFFER option */
+ if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp))
+ return -EIO;
+ opt_size_pos = ftell(trace_dat_fp);
+ if (!fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp))
+ return -EIO;
+ opt_payload_start = ftell(trace_dat_fp);
+
+ /* data_offset placeholder */
+ if (!fwrite(&data_offset, sizeof(unsigned long long), 1, trace_dat_fp) ||
+ !fwrite("\0", 1, 1, trace_dat_fp) ||
+ !fwrite(clock, 1, strlen(clock) + 1, trace_dat_fp) ||
+ !fwrite(&page_size, sizeof(unsigned int), 1, trace_dat_fp) ||
+ !fwrite(&trace_dat_nr_cpus, sizeof(unsigned int), 1, trace_dat_fp))
+ return -EIO;
+
+ /* per cpu: cpu_id + offset placeholder + size */
+ for (cpu = 0; cpu < trace_dat_nr_cpus; cpu++) {
+ cpu_offset = 0; /* filled in write_flyrecord */
+ cpu_size = 0; /* filled in write_flyrecord */
+
+ if (!fwrite(&cpu, sizeof(unsigned int), 1, trace_dat_fp))
+ return -EIO;
+ buffer_opt_cpu_offsets_pos[cpu] = ftell(trace_dat_fp);
+ if (!fwrite(&cpu_offset, sizeof(unsigned long long), 1, trace_dat_fp) ||
+ !fwrite(&cpu_size, sizeof(unsigned long long), 1, trace_dat_fp))
+ return -EIO;
+ }
+
+ /* fill opt_size */
+ end_pos = ftell(trace_dat_fp);
+
+ opt_size = end_pos - opt_payload_start;
+ fseek(trace_dat_fp, opt_size_pos, SEEK_SET);
+ if (!fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp))
+ return -EIO;
+ fseek(trace_dat_fp, end_pos, SEEK_SET);
+
+ /* DONE id=0 */
+ done_id = TRACE_DAT_OPTION_DONE;
+ done_size = sizeof(unsigned long long);
+ /* No additional options sections follow this one */
+ next = 0;
+
+ if (!fwrite(&done_id, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&done_size, sizeof(unsigned int), 1, trace_dat_fp) ||
+ !fwrite(&next, sizeof(unsigned long long), 1, trace_dat_fp))
+ return -EIO;
+
+ /* fill section size */
+ end_pos = ftell(trace_dat_fp);
+
+ section_size = end_pos - payload_start;
+ fseek(trace_dat_fp, section_size_pos, SEEK_SET);
+ if (!fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp))
+ return -EIO;
+ fseek(trace_dat_fp, end_pos, SEEK_SET);
+
+ return 0;
+
+}
+
+int trace_dat__write_flyrecord_section(void)
+{
+ unsigned short section_id = TRACE_DAT_SECTION_FLYRECORD;
+ unsigned short flags = 0;
+ unsigned int string_id = STRID_BUFFER_FLYRECORD;
+ unsigned long long section_size = 0;
+ long section_size_pos;
+ long flyrecord_start;
+ long after_header;
+ long padding_needed;
+ unsigned long long *cpu_offsets;
+ unsigned long long *cpu_sizes;
+ int cpu;
+ int ret = 0;
+ char *pad;
+ unsigned long long start;
+ long end_pos;
+
+ if (!trace_dat_fp)
+ return -EINVAL;
+
+ cpu_offsets = calloc(trace_dat_nr_cpus, sizeof(unsigned long long));
+ cpu_sizes = calloc(trace_dat_nr_cpus, sizeof(unsigned long long));
+ if (!cpu_offsets || !cpu_sizes) {
+ ret = -ENOMEM;
+ goto cleanup;
+ }
+ flyrecord_start = ftell(trace_dat_fp);
+ if (flyrecord_start < 0) {
+ ret = -EIO;
+ goto cleanup;
+ }
+
+ /* section header */
+ if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp)) {
+ ret = -EIO;
+ goto cleanup;
+ }
+ section_size_pos = ftell(trace_dat_fp);
+ if (!fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp)) {
+ ret = -EIO;
+ goto cleanup;
+ }
+
+ /* Align to page boundary */
+ after_header = ftell(trace_dat_fp);
+ padding_needed = (trace_dat_page_size -
+ (after_header % trace_dat_page_size)) % trace_dat_page_size;
+
+ if (padding_needed > 0) {
+ pad = calloc(1, padding_needed);
+ if (!pad) {
+ ret = -ENOMEM;
+ goto cleanup;
+ }
+
+ if (!fwrite(pad, 1, padding_needed, trace_dat_fp)) {
+ free(pad);
+ ret = -EIO;
+ goto cleanup;
+ }
+ free(pad);
+ }
+
+ /* write per-cpu trace data */
+ for (cpu = 0; cpu < trace_dat_nr_cpus; cpu++) {
+ start = ftell(trace_dat_fp);
+
+ ret = trace_dat__write_cpu_dat(trace_dat_fp, cpu, &cpu_offsets[cpu]);
+
+ if (ret < 0) {
+ pr_err("Failed to write CPU %d data\n", cpu);
+ goto cleanup;
+ }
+ cpu_sizes[cpu] = ftell(trace_dat_fp) - start;
+ }
+
+ /* fill section size */
+ end_pos = ftell(trace_dat_fp);
+
+ section_size = end_pos - after_header;
+ if (fseek(trace_dat_fp, section_size_pos, SEEK_SET) < 0 ||
+ !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp)) {
+ ret = -EIO;
+ goto cleanup;
+ }
+ if (fseek(trace_dat_fp, end_pos, SEEK_SET) < 0) {
+ ret = -EIO;
+ goto cleanup;
+ }
+
+ /* fill cpu offsets and sizes in BUFFER option */
+ for (cpu = 0; cpu < trace_dat_nr_cpus; cpu++) {
+ if (fseek(trace_dat_fp, buffer_opt_cpu_offsets_pos[cpu], SEEK_SET) < 0 ||
+ !fwrite(&cpu_offsets[cpu], sizeof(unsigned long long), 1, trace_dat_fp) ||
+ !fwrite(&cpu_sizes[cpu], sizeof(unsigned long long), 1, trace_dat_fp)) {
+ ret = -EIO;
+ goto cleanup;
+ }
+ }
+
+ /* fill data offset in buffer option */
+ if (fseek(trace_dat_fp, opt_payload_start, SEEK_SET) < 0 ||
+ !fwrite(&flyrecord_start, sizeof(unsigned long long), 1, trace_dat_fp)) {
+ ret = -EIO;
+ goto cleanup;
+ }
+
+ if (fseek(trace_dat_fp, 0, SEEK_END) < 0) {
+ ret = -EIO;
+ goto cleanup;
+ }
+
+
+cleanup:
+ free(cpu_offsets);
+ free(cpu_sizes);
+ return ret;
+}
+
+/* Free all per-CPU event buffers */
+void trace_dat__free_cpu_buffers(void)
+{
+ int cpu;
+
+ if (!trace_cpu_data)
+ return;
+
+ for (cpu = 0; cpu < trace_dat_nr_cpus; cpu++) {
+ int i;
+
+ for (i = 0; i < trace_cpu_data[cpu].count; i++)
+ free(trace_cpu_data[cpu].events[i].raw);
+ free(trace_cpu_data[cpu].events);
+ }
+ free(trace_cpu_data);
+ trace_cpu_data = NULL;
+ free(buffer_opt_cpu_offsets_pos);
+ buffer_opt_cpu_offsets_pos = NULL;
+ trace_dat_nr_cpus = 0;
+}
diff --git a/tools/perf/util/trace-dat.h b/tools/perf/util/trace-dat.h
new file mode 100644
index 000000000000..9aec37b708d4
--- /dev/null
+++ b/tools/perf/util/trace-dat.h
@@ -0,0 +1,83 @@
+/* SPDX-License-Identifier: GPL-2.0-or-later */
+/*
+ * Copyright 2026, IBM Corporation
+ * Author: Tanushree Shah <tshah@linux.ibm.com>
+ */
+
+#ifndef __PERF_TRACE_DAT_H
+#define __PERF_TRACE_DAT_H
+
+#include <stdio.h>
+
+/* trace.dat file format version */
+#define TRACE_DAT_VERSION '7'
+
+/*
+ * Section IDs for trace.dat format
+ */
+#define TRACE_DAT_SECTION_OPTIONS 0
+#define TRACE_DAT_SECTION_FLYRECORD 3
+#define TRACE_DAT_SECTION_STRINGS 15
+#define TRACE_DAT_SECTION_HEADER 16
+#define TRACE_DAT_SECTION_FTRACE 17
+#define TRACE_DAT_SECTION_EVENTS 18
+#define TRACE_DAT_SECTION_KALLSYMS 19
+#define TRACE_DAT_SECTION_CMDLINE 21
+
+/*
+ * Option IDs for trace.dat options sections
+ */
+#define TRACE_DAT_OPTION_DONE 0
+#define TRACE_DAT_OPTION_BUFFER 3
+#define TRACE_DAT_OPTION_TRACECLOCK 4
+#define TRACE_DAT_OPTION_CPUCOUNT 8
+#define TRACE_DAT_OPTION_HEADER 16
+#define TRACE_DAT_OPTION_FTRACE 17
+#define TRACE_DAT_OPTION_EVENT 18
+#define TRACE_DAT_OPTION_KALLSYMS 19
+#define TRACE_DAT_OPTION_CMDLINE 21
+
+/*
+ * String offsets in the strings section
+ * These point to null-terminated strings used as section names
+ */
+#define STRID_HEADERS 0
+#define STRID_FTRACE_FORMATS 8
+#define STRID_EVENT_FORMATS 29
+#define STRID_KALLSYMS 43
+#define STRID_CMDLINES 52
+#define STRID_STRINGS 61
+#define STRID_OPTIONS_1 69
+#define STRID_OPTIONS_2 77
+#define STRID_BUFFER_FLYRECORD 85
+
+struct perf_session;
+
+extern FILE *trace_dat_fp;
+extern int trace_dat_page_size;
+extern int trace_dat_nr_cpus;
+extern long trace_dat_options_offset;
+extern long trace_dat_header_info_offset;
+extern long trace_dat_events_format_offset;
+extern long trace_dat_ftrace_format_offset;
+extern long trace_dat_kallsyms_offset;
+extern long trace_dat_cmdline_offset;
+extern long trace_dat_next_options_offset;
+
+/* collect and manage per-cpu tracepoint event buffers */
+#define TRACE_DAT_RECORD_TYPE_TIME_EXTEND 30
+#define TRACE_DAT_RECORD_TIME_EXTEND_SIZE 8
+#define TRACE_DAT_RECORD_TYPE_LEN_MAX 28
+
+int trace_dat__init_cpu_buffers(int nr_cpus);
+int trace_dat__collect_cpu_event(int cpu, unsigned long long ts,
+ void *raw, unsigned int raw_size);
+void trace_dat__free_cpu_buffers(void);
+
+/* write trace.dat file sections */
+int trace_dat__write_options_section1(void);
+int trace_dat__write_options_section2(void);
+int trace_dat__write_flyrecord_section(void);
+int trace_dat__write_strings_section(void);
+
+#endif /* __PERF_TRACE_DAT_H */
--
2.47.1
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [RFC PATCH v4 2/5] perf/trace-event: Write trace.dat metadata sections during parsing
2026-08-22 6:27 [RFC PATCH v4 0/5] Add perf.data tracepoint events to trace.dat conversion Tanushree Shah
2026-08-22 6:27 ` [RFC PATCH v4 1/5] perf/trace-dat: Add trace.dat export infrastructure Tanushree Shah
@ 2026-08-22 6:27 ` Tanushree Shah
2026-08-22 6:44 ` sashiko-bot
2026-08-22 6:27 ` [RFC PATCH v4 3/5] perf data-convert: Add perf.data to trace.dat conversion backend Tanushree Shah
` (2 subsequent siblings)
4 siblings, 1 reply; 11+ messages in thread
From: Tanushree Shah @ 2026-08-22 6:27 UTC (permalink / raw)
To: acme, jolsa, adrian.hunter, vmolnaro, mpetlan, tmricht, maddy,
irogers, namhyung
Cc: linux-perf-users, linuxppc-dev, atrajeev, hbathini, Tejas.Manhas1,
Tanushree.Shah, rostedt, Tanushree Shah
Perf already captures the tracing metadata as a part of
data section in perf.data
When trace_dat_fp is set, write trace.dat compatible metadata
sections using the perf provided raw buffers.
Sections written:
- Initial format header (magic, version, endian, long_size,
page_size, compression, options_offset placeholder)
- Section 16: HEADER INFO (header_page + header_event)
- Section 17: FTRACE EVENT FORMATS
- Section 18: EVENT FORMATS (per system/event format files)
- Section 19: KALLSYMS
- Section 21: CMDLINES
- Section 15: STRINGS (written last after all sections)
Introduce to_file_u16/u32/u64 helpers wrapping tep_read_number() to
write all trace.dat metadata in the recorded machine's byte
order, ensuring correct cross-platform compatibility.
Signed-off-by: Tanushree Shah <tshah@linux.ibm.com>
---
tools/perf/util/trace-dat.c | 212 ++++++++++++--------
tools/perf/util/trace-dat.h | 40 +++-
tools/perf/util/trace-event-read.c | 307 ++++++++++++++++++++++++++++-
3 files changed, 465 insertions(+), 94 deletions(-)
diff --git a/tools/perf/util/trace-dat.c b/tools/perf/util/trace-dat.c
index 55a7bd982c6f..f71e03716e27 100644
--- a/tools/perf/util/trace-dat.c
+++ b/tools/perf/util/trace-dat.c
@@ -43,6 +43,7 @@
/* Buffer size for reading trace_clock string from debugfs/tracefs */
#define CLOCK_BUFFER_SIZE 256
+bool trace_dat_write_failed;
FILE *trace_dat_fp;
int trace_dat_page_size;
int trace_dat_nr_cpus;
@@ -143,10 +144,12 @@ int trace_dat__collect_cpu_event(int cpu, unsigned long long ts,
}
/* Write a single page of trace records */
-static int trace_dat__write_page(FILE *fp, unsigned long long base_ts,
+static int trace_dat__write_page(FILE *fp, struct tep_handle *pevent, unsigned long long base_ts,
char **records, int *rec_sizes, int nr_recs)
{
unsigned long long commit = 0;
+ unsigned long long ts_out;
+ unsigned long long commit_out;
int offset = TRACE_DAT_RECORD_HEADER_SIZE;
int i;
char *page;
@@ -161,8 +164,12 @@ static int trace_dat__write_page(FILE *fp, unsigned long long base_ts,
commit += rec_sizes[i];
}
- memcpy(page, &base_ts, sizeof(base_ts));
- memcpy(page + sizeof(base_ts), &commit, sizeof(commit));
+ /* Byte-swap page header for cross-arch compatibility */
+ ts_out = to_file_u64(pevent, base_ts);
+ commit_out = to_file_u64(pevent, commit);
+
+ memcpy(page, &ts_out, sizeof(ts_out));
+ memcpy(page + sizeof(ts_out), &commit_out, sizeof(commit_out));
if (!fwrite(page, 1, trace_dat_page_size, fp)) {
free(page);
@@ -174,7 +181,8 @@ static int trace_dat__write_page(FILE *fp, unsigned long long base_ts,
}
/* Write all trace data for a single CPU as trace.dat flyrecord pages */
-static int trace_dat__write_cpu_dat(FILE *fp, int cpu, unsigned long long *file_offset_out)
+static int trace_dat__write_cpu_dat(FILE *fp, struct tep_handle *pevent,
+ int cpu, unsigned long long *file_offset_out)
{
struct cpu_events *cpu_events = &trace_cpu_data[cpu];
unsigned long long base_ts;
@@ -229,11 +237,18 @@ static int trace_dat__write_cpu_dat(FILE *fp, int cpu, unsigned long long *file_
if (!extend)
return -ENOMEM;
- extend_hdr =
- ((time_delta & TRACE_DAT_RECORD_TIME_MASK) <<
- TRACE_DAT_RECORD_TIME_SHIFT) |
- TRACE_DAT_RECORD_TYPE_TIME_EXTEND;
- delta_upper = time_delta >> TRACE_DAT_RECORD_TIME_SHIFT;
+ if (tep_is_file_bigendian(pevent)) {
+ extend_hdr = (time_delta & TRACE_DAT_RECORD_TIME_MASK) |
+ (TRACE_DAT_RECORD_TYPE_TIME_EXTEND << 27);
+ delta_upper = time_delta >> 27;
+ } else {
+ extend_hdr = ((time_delta & TRACE_DAT_RECORD_TIME_MASK) <<
+ TRACE_DAT_RECORD_TIME_SHIFT) |
+ TRACE_DAT_RECORD_TYPE_TIME_EXTEND;
+ delta_upper = time_delta >> TRACE_DAT_RECORD_TIME_SHIFT;
+ }
+ extend_hdr = to_file_u32(pevent, extend_hdr); /* still needed */
+ delta_upper = to_file_u32(pevent, delta_upper); /* still needed */
memcpy(extend, &extend_hdr, TRACE_DAT_WORD_SIZE);
memcpy(extend + TRACE_DAT_WORD_SIZE, &delta_upper,
@@ -273,7 +288,7 @@ static int trace_dat__write_cpu_dat(FILE *fp, int cpu, unsigned long long *file_
/* Check page fit BEFORE allocating data record */
if (page_size_used + needed_size >
trace_dat_page_size - TRACE_DAT_RECORD_HEADER_SIZE) {
- ret = trace_dat__write_page(fp, base_ts,
+ ret = trace_dat__write_page(fp, pevent, base_ts,
page_records, page_rec_sizes,
nr_page_recs);
@@ -299,7 +314,13 @@ static int trace_dat__write_cpu_dat(FILE *fp, int cpu, unsigned long long *file_
time_delta = 0;
}
- hdr_word = (time_delta << TRACE_DAT_RECORD_TIME_SHIFT) | type_len;
+ if (tep_is_file_bigendian(pevent))
+ hdr_word = ((unsigned int)time_delta & TRACE_DAT_RECORD_TIME_MASK) |
+ (type_len << 27);
+ else
+ hdr_word = ((unsigned int)time_delta << TRACE_DAT_RECORD_TIME_SHIFT) |
+ type_len;
+ hdr_word = to_file_u32(pevent, hdr_word);
data_rec = calloc(1, data_rec_size);
if (!data_rec) {
@@ -311,8 +332,11 @@ static int trace_dat__write_cpu_dat(FILE *fp, int cpu, unsigned long long *file_
memcpy(data_rec, &hdr_word, TRACE_DAT_WORD_SIZE);
/* Large events: write actual byte length after header */
- if (type_len == 0)
- memcpy(data_rec + TRACE_DAT_WORD_SIZE, &data_len, TRACE_DAT_WORD_SIZE);
+ if (type_len == 0) {
+ unsigned int data_len_out = to_file_u32(pevent, data_len);
+
+ memcpy(data_rec + TRACE_DAT_WORD_SIZE, &data_len_out, TRACE_DAT_WORD_SIZE);
+ }
memcpy(data_rec + payload_offset, event->raw, data_len);
@@ -366,7 +390,7 @@ static int trace_dat__write_cpu_dat(FILE *fp, int cpu, unsigned long long *file_
}
if (nr_page_recs > 0) {
- ret = trace_dat__write_page(fp, base_ts,
+ ret = trace_dat__write_page(fp, pevent, base_ts,
page_records, page_rec_sizes, nr_page_recs);
}
out_free:
@@ -378,11 +402,11 @@ static int trace_dat__write_cpu_dat(FILE *fp, int cpu, unsigned long long *file_
}
/* Write the strings section containing section name lookup table */
-int trace_dat__write_strings_section(void)
+int trace_dat__write_strings_section(struct tep_handle *pevent)
{
- unsigned short section_id = TRACE_DAT_SECTION_STRINGS;
- unsigned short flags = 0;
unsigned long long section_size = 0;
+ unsigned short section_id = to_file_u16(pevent, TRACE_DAT_SECTION_STRINGS);
+ unsigned short flags = to_file_u16(pevent, 0);
static const char * const section_names[] = {
"headers", /* offset 0 - strid for section 16 */
"ftrace event formats", /* offset 8 - strid for section 17 */
@@ -397,7 +421,7 @@ int trace_dat__write_strings_section(void)
};
/* string_id points to "strings" string itself */
- unsigned int string_id = STRID_STRINGS;
+ unsigned int string_id = to_file_u32(pevent, STRID_STRINGS);
int i;
if (!trace_dat_fp)
@@ -405,47 +429,59 @@ int trace_dat__write_strings_section(void)
for (i = 0; section_names[i] != NULL; i++)
section_size += strlen(section_names[i]) + 1;
+ section_size = to_file_u64(pevent, section_size);
/* write section header */
if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) ||
!fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) ||
!fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp) ||
- !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp))
- return -EIO;
+ !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp)) {
+ pr_warning("Failed to write strings section\n");
+ trace_dat_write_failed = true;
+ }
/* write strings */
for (i = 0; section_names[i] != NULL; i++)
- if (!fwrite(section_names[i], 1, strlen(section_names[i]) + 1, trace_dat_fp))
- return -EIO;
+ if (!fwrite(section_names[i], 1, strlen(section_names[i]) + 1, trace_dat_fp)) {
+ pr_warning("Failed to write strings section\n");
+ trace_dat_write_failed = true;
+ }
return 0;
}
/* Writes options section containing CPUCOUNT, TRACECLOCK, EVENT_FORMAT, HEADER_INFO,
* FTRACE_EVENTS, KALLSYMS, CMDLINES options, ending with DONE option pointing to next section.
*/
-int trace_dat__write_options_section1(void)
+int trace_dat__write_options_section1(struct tep_handle *pevent)
{
- unsigned short section_id = TRACE_DAT_SECTION_OPTIONS;
- unsigned short flags = 0;
- unsigned int string_id = STRID_OPTIONS_1;
+ unsigned short section_id = to_file_u16(pevent, TRACE_DAT_SECTION_OPTIONS);
+ unsigned short flags = to_file_u16(pevent, 0);
+ unsigned int string_id = to_file_u32(pevent, STRID_OPTIONS_1);
unsigned long long section_size = 0;
long section_size_pos;
long payload_start;
unsigned long long section_start;
unsigned short opt_id;
unsigned int opt_size;
+ unsigned int raw_opt_size;
+ unsigned int nr_cpus;
char clock_buf[CLOCK_BUFFER_SIZE];
FILE *clock_file;
size_t bytes_read;
char *path;
unsigned long long next_offset;
+ unsigned long long events_off;
+ unsigned long long header_off;
+ unsigned long long ftrace_off;
+ unsigned long long kallsyms_off;
+ unsigned long long cmdline_off;
long end_pos;
if (!trace_dat_fp)
return -EBADF;
/* fill options_offset in initial format */
- section_start = ftell(trace_dat_fp);
+ section_start = to_file_u64(pevent, ftell(trace_dat_fp));
if (fseek(trace_dat_fp, trace_dat_options_offset, SEEK_SET) < 0 ||
!fwrite(§ion_start, sizeof(unsigned long long), 1, trace_dat_fp) ||
@@ -464,16 +500,17 @@ int trace_dat__write_options_section1(void)
payload_start = ftell(trace_dat_fp);
/* CPUCOUNT option */
- opt_id = TRACE_DAT_OPTION_CPUCOUNT;
- opt_size = sizeof(unsigned int);
+ opt_id = to_file_u16(pevent, TRACE_DAT_OPTION_CPUCOUNT);
+ opt_size = to_file_u32(pevent, sizeof(unsigned int));
+ nr_cpus = to_file_u32(pevent, trace_dat_nr_cpus);
if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) ||
!fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) ||
- !fwrite(&trace_dat_nr_cpus, sizeof(unsigned int), 1, trace_dat_fp))
+ !fwrite(&nr_cpus, sizeof(unsigned int), 1, trace_dat_fp))
return -EIO;
/* TRACECLOCK option */
- opt_id = TRACE_DAT_OPTION_TRACECLOCK;
+ opt_id = to_file_u16(pevent, TRACE_DAT_OPTION_TRACECLOCK);
path = get_tracing_file("trace_clock");
if (path) {
@@ -490,66 +527,72 @@ int trace_dat__write_options_section1(void)
strcpy(clock_buf, "local\n");
bytes_read = strlen(clock_buf);
}
- opt_size = bytes_read + 1;
+ raw_opt_size = bytes_read + 1;
+ opt_size = to_file_u32(pevent, bytes_read + 1);
if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) ||
!fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) ||
- !fwrite(clock_buf, 1, opt_size, trace_dat_fp))
+ !fwrite(clock_buf, 1, raw_opt_size, trace_dat_fp))
return -EIO;
/* EVENT option */
- opt_id = TRACE_DAT_OPTION_EVENT;
- opt_size = sizeof(unsigned long long);
+ opt_id = to_file_u16(pevent, TRACE_DAT_OPTION_EVENT);
+ opt_size = to_file_u32(pevent, sizeof(unsigned long long));
+ events_off = to_file_u64(pevent, trace_dat_events_format_offset);
if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) ||
!fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) ||
- !fwrite(&trace_dat_events_format_offset, sizeof(unsigned long long),
+ !fwrite(&events_off, sizeof(unsigned long long),
1, trace_dat_fp))
return -EIO;
/* HEADER option */
- opt_id = TRACE_DAT_OPTION_HEADER;
- opt_size = sizeof(unsigned long long);
+ opt_id = to_file_u16(pevent, TRACE_DAT_OPTION_HEADER);
+ opt_size = to_file_u32(pevent, sizeof(unsigned long long));
+ header_off = to_file_u64(pevent, trace_dat_header_info_offset);
if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) ||
!fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) ||
- !fwrite(&trace_dat_header_info_offset, sizeof(unsigned long long),
+ !fwrite(&header_off, sizeof(unsigned long long),
1, trace_dat_fp))
return -EIO;
/* FTRACE option */
- opt_id = TRACE_DAT_OPTION_FTRACE;
- opt_size = sizeof(unsigned long long);
+ opt_id = to_file_u16(pevent, TRACE_DAT_OPTION_FTRACE);
+ opt_size = to_file_u32(pevent, sizeof(unsigned long long));
+ ftrace_off = to_file_u64(pevent, trace_dat_ftrace_format_offset);
if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) ||
!fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) ||
- !fwrite(&trace_dat_ftrace_format_offset, sizeof(unsigned long long),
+ !fwrite(&ftrace_off, sizeof(unsigned long long),
1, trace_dat_fp))
return -EIO;
/* KALLSYMS option */
- opt_id = TRACE_DAT_OPTION_KALLSYMS;
- opt_size = sizeof(unsigned long long);
+ opt_id = to_file_u16(pevent, TRACE_DAT_OPTION_KALLSYMS);
+ opt_size = to_file_u32(pevent, sizeof(unsigned long long));
+ kallsyms_off = to_file_u64(pevent, trace_dat_kallsyms_offset);
if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) ||
!fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) ||
- !fwrite(&trace_dat_kallsyms_offset, sizeof(unsigned long long),
+ !fwrite(&kallsyms_off, sizeof(unsigned long long),
1, trace_dat_fp))
return -EIO;
/* CMDLINE option */
- opt_id = TRACE_DAT_OPTION_CMDLINE;
- opt_size = sizeof(unsigned long long);
+ opt_id = to_file_u16(pevent, TRACE_DAT_OPTION_CMDLINE);
+ opt_size = to_file_u32(pevent, sizeof(unsigned long long));
+ cmdline_off = to_file_u64(pevent, trace_dat_cmdline_offset);
if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) ||
!fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) ||
- !fwrite(&trace_dat_cmdline_offset, sizeof(unsigned long long),
+ !fwrite(&cmdline_off, sizeof(unsigned long long),
1, trace_dat_fp))
return -EIO;
/* DONE option id - next_options_offset filled later */
- opt_id = TRACE_DAT_OPTION_DONE;
- opt_size = sizeof(unsigned long long);
- next_offset = 0; /* placeholder */
+ opt_id = to_file_u16(pevent, TRACE_DAT_OPTION_DONE);
+ opt_size = to_file_u32(pevent, sizeof(unsigned long long));
+ next_offset = to_file_u64(pevent, 0);
trace_dat_next_options_offset = ftell(trace_dat_fp);
if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) ||
@@ -560,7 +603,7 @@ int trace_dat__write_options_section1(void)
/* fill section size */
end_pos = ftell(trace_dat_fp);
- section_size = end_pos - payload_start;
+ section_size = to_file_u64(pevent, end_pos - payload_start);
if (fseek(trace_dat_fp, section_size_pos, SEEK_SET) < 0 ||
!fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp) ||
fseek(trace_dat_fp, end_pos, SEEK_SET) < 0)
@@ -574,20 +617,21 @@ int trace_dat__write_options_section1(void)
* (flyrecord section offset, clock type, page size, CPU count,
* per-CPU offsets/sizes) and DONE option.
*/
-int trace_dat__write_options_section2(void)
+int trace_dat__write_options_section2(struct tep_handle *pevent)
{
- unsigned short section_id = TRACE_DAT_SECTION_OPTIONS;
- unsigned short flags = 0;
- unsigned int string_id = STRID_OPTIONS_2;
+ unsigned short section_id = to_file_u16(pevent, TRACE_DAT_SECTION_OPTIONS);
+ unsigned short flags = to_file_u16(pevent, 0);
+ unsigned int string_id = to_file_u32(pevent, STRID_OPTIONS_2);
unsigned long long section_size = 0;
long section_size_pos;
long payload_start;
int cpu;
- unsigned short opt_id = TRACE_DAT_OPTION_BUFFER;
- unsigned int opt_size = 0;
+ unsigned short opt_id = to_file_u16(pevent, TRACE_DAT_OPTION_BUFFER);
+ unsigned int opt_size = 0; /* filled after payload written */
long opt_size_pos;
unsigned long long data_offset = 0;
- unsigned int page_size = (unsigned int)trace_dat_page_size;
+ unsigned int page_size = to_file_u32(pevent, (unsigned int)trace_dat_page_size);
+ unsigned int nr_cpus = to_file_u32(pevent, trace_dat_nr_cpus);
const char *clock = "local";
unsigned long long next;
long end_pos;
@@ -600,7 +644,7 @@ int trace_dat__write_options_section2(void)
return -EINVAL;
/* fill done1 next offset - points to this section */
- next = ftell(trace_dat_fp);
+ next = to_file_u64(pevent, ftell(trace_dat_fp));
if (fseek(trace_dat_fp, trace_dat_next_options_offset + 2 + 4, SEEK_SET) < 0 ||
!fwrite(&next, sizeof(unsigned long long), 1, trace_dat_fp) ||
@@ -631,15 +675,16 @@ int trace_dat__write_options_section2(void)
!fwrite("\0", 1, 1, trace_dat_fp) ||
!fwrite(clock, 1, strlen(clock) + 1, trace_dat_fp) ||
!fwrite(&page_size, sizeof(unsigned int), 1, trace_dat_fp) ||
- !fwrite(&trace_dat_nr_cpus, sizeof(unsigned int), 1, trace_dat_fp))
+ !fwrite(&nr_cpus, sizeof(unsigned int), 1, trace_dat_fp))
return -EIO;
/* per cpu: cpu_id + offset placeholder + size */
for (cpu = 0; cpu < trace_dat_nr_cpus; cpu++) {
+ unsigned int cpu_out = to_file_u32(pevent, cpu);
cpu_offset = 0; /* filled in write_flyrecord */
- cpu_size = 0; /* filled in write_flyrecord */
+ cpu_size = 0; /* filled in write_flyrecord */
- if (!fwrite(&cpu, sizeof(unsigned int), 1, trace_dat_fp))
+ if (!fwrite(&cpu_out, sizeof(unsigned int), 1, trace_dat_fp))
return -EIO;
buffer_opt_cpu_offsets_pos[cpu] = ftell(trace_dat_fp);
if (!fwrite(&cpu_offset, sizeof(unsigned long long), 1, trace_dat_fp) ||
@@ -650,17 +695,17 @@ int trace_dat__write_options_section2(void)
/* fill opt_size */
end_pos = ftell(trace_dat_fp);
- opt_size = end_pos - opt_payload_start;
- fseek(trace_dat_fp, opt_size_pos, SEEK_SET);
- if (!fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp))
+ opt_size = to_file_u32(pevent, end_pos - opt_payload_start);
+ if (fseek(trace_dat_fp, opt_size_pos, SEEK_SET) < 0 ||
+ !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) ||
+ fseek(trace_dat_fp, end_pos, SEEK_SET) < 0)
return -EIO;
- fseek(trace_dat_fp, end_pos, SEEK_SET);
/* DONE id=0 */
- done_id = TRACE_DAT_OPTION_DONE;
- done_size = sizeof(unsigned long long);
+ done_id = to_file_u16(pevent, TRACE_DAT_OPTION_DONE);
+ done_size = to_file_u32(pevent, sizeof(unsigned long long));
/* No additional options sections follow this one */
- next = 0;
+ next = to_file_u64(pevent, 0);
if (!fwrite(&done_id, sizeof(unsigned short), 1, trace_dat_fp) ||
!fwrite(&done_size, sizeof(unsigned int), 1, trace_dat_fp) ||
@@ -670,24 +715,25 @@ int trace_dat__write_options_section2(void)
/* fill section size */
end_pos = ftell(trace_dat_fp);
- section_size = end_pos - payload_start;
- fseek(trace_dat_fp, section_size_pos, SEEK_SET);
- if (!fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp))
+ section_size = to_file_u64(pevent, end_pos - payload_start);
+ if (fseek(trace_dat_fp, section_size_pos, SEEK_SET) < 0 ||
+ !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp) ||
+ fseek(trace_dat_fp, end_pos, SEEK_SET) < 0)
return -EIO;
- fseek(trace_dat_fp, end_pos, SEEK_SET);
return 0;
}
-int trace_dat__write_flyrecord_section(void)
+int trace_dat__write_flyrecord_section(struct tep_handle *pevent)
{
- unsigned short section_id = TRACE_DAT_SECTION_FLYRECORD;
- unsigned short flags = 0;
- unsigned int string_id = STRID_BUFFER_FLYRECORD;
+ unsigned short section_id = to_file_u16(pevent, TRACE_DAT_SECTION_FLYRECORD);
+ unsigned short flags = to_file_u16(pevent, 0);
+ unsigned int string_id = to_file_u32(pevent, STRID_BUFFER_FLYRECORD);
unsigned long long section_size = 0;
long section_size_pos;
long flyrecord_start;
+ unsigned long long flyrecord_out;
long after_header;
long padding_needed;
unsigned long long *cpu_offsets;
@@ -750,7 +796,7 @@ int trace_dat__write_flyrecord_section(void)
for (cpu = 0; cpu < trace_dat_nr_cpus; cpu++) {
start = ftell(trace_dat_fp);
- ret = trace_dat__write_cpu_dat(trace_dat_fp, cpu, &cpu_offsets[cpu]);
+ ret = trace_dat__write_cpu_dat(trace_dat_fp, pevent, cpu, &cpu_offsets[cpu]);
if (ret < 0) {
pr_err("Failed to write CPU %d data\n", cpu);
@@ -762,7 +808,7 @@ int trace_dat__write_flyrecord_section(void)
/* fill section size */
end_pos = ftell(trace_dat_fp);
- section_size = end_pos - after_header;
+ section_size = to_file_u64(pevent, end_pos - after_header);
if (fseek(trace_dat_fp, section_size_pos, SEEK_SET) < 0 ||
!fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp)) {
ret = -EIO;
@@ -775,6 +821,8 @@ int trace_dat__write_flyrecord_section(void)
/* fill cpu offsets and sizes in BUFFER option */
for (cpu = 0; cpu < trace_dat_nr_cpus; cpu++) {
+ cpu_offsets[cpu] = to_file_u64(pevent, cpu_offsets[cpu]);
+ cpu_sizes[cpu] = to_file_u64(pevent, cpu_sizes[cpu]);
if (fseek(trace_dat_fp, buffer_opt_cpu_offsets_pos[cpu], SEEK_SET) < 0 ||
!fwrite(&cpu_offsets[cpu], sizeof(unsigned long long), 1, trace_dat_fp) ||
!fwrite(&cpu_sizes[cpu], sizeof(unsigned long long), 1, trace_dat_fp)) {
@@ -784,8 +832,9 @@ int trace_dat__write_flyrecord_section(void)
}
/* fill data offset in buffer option */
+ flyrecord_out = to_file_u64(pevent, flyrecord_start);
if (fseek(trace_dat_fp, opt_payload_start, SEEK_SET) < 0 ||
- !fwrite(&flyrecord_start, sizeof(unsigned long long), 1, trace_dat_fp)) {
+ !fwrite(&flyrecord_out, sizeof(unsigned long long), 1, trace_dat_fp)) {
ret = -EIO;
goto cleanup;
}
@@ -795,7 +844,6 @@ int trace_dat__write_flyrecord_section(void)
goto cleanup;
}
-
cleanup:
free(cpu_offsets);
free(cpu_sizes);
diff --git a/tools/perf/util/trace-dat.h b/tools/perf/util/trace-dat.h
index 9aec37b708d4..5985083b275a 100644
--- a/tools/perf/util/trace-dat.h
+++ b/tools/perf/util/trace-dat.h
@@ -8,9 +8,13 @@
#define __PERF_TRACE_DAT_H
#include <stdio.h>
+#include <stdbool.h>
+#include <event-parse.h>
+#include <byteswap.h>
+#include "util.h"
/* trace.dat file format version */
-#define TRACE_DAT_VERSION '7'
+#define TRACE_DAT_VERSION "7"
/*
* Section IDs for trace.dat format
@@ -51,6 +55,32 @@
#define STRID_OPTIONS_2 77
#define STRID_BUFFER_FLYRECORD 85
+/*
+ * Set by trace.dat section writers (kallsyms, ftrace, events etc.) on
+ * fwrite() failure; checked by trace_convert__perf2dat() to clean up.
+ */
+extern bool trace_dat_write_failed;
+
+/*
+ * tep handle from the recording machine; set by trace_convert__perf2dat()
+ * before options/flyrecord sections are written.
+ */
+
+static inline uint16_t to_file_u16(struct tep_handle *pevent, uint16_t val)
+{
+ return tep_read_number(pevent, &val, 2);
+}
+
+static inline uint32_t to_file_u32(struct tep_handle *pevent, uint32_t val)
+{
+ return tep_read_number(pevent, &val, 4);
+}
+
+static inline uint64_t to_file_u64(struct tep_handle *pevent, uint64_t val)
+{
+ return tep_read_number(pevent, &val, 8);
+}
+
struct perf_session;
extern FILE *trace_dat_fp;
@@ -75,9 +105,9 @@ int trace_dat__collect_cpu_event(int cpu, unsigned long long ts,
void trace_dat__free_cpu_buffers(void);
/* write trace.dat file sections */
-int trace_dat__write_options_section1(void);
-int trace_dat__write_options_section2(void);
-int trace_dat__write_flyrecord_section(void);
-int trace_dat__write_strings_section(void);
+int trace_dat__write_options_section1(struct tep_handle *pevent);
+int trace_dat__write_options_section2(struct tep_handle *pevent);
+int trace_dat__write_flyrecord_section(struct tep_handle *pevent);
+int trace_dat__write_strings_section(struct tep_handle *pevent);
#endif /* __PERF_TRACE_DAT_H */
diff --git a/tools/perf/util/trace-event-read.c b/tools/perf/util/trace-event-read.c
index db1622fc99c2..dd81c675e43c 100644
--- a/tools/perf/util/trace-event-read.c
+++ b/tools/perf/util/trace-event-read.c
@@ -20,6 +20,7 @@
#include "trace-event.h"
#include "debug.h"
#include "util.h"
+#include "trace-dat.h"
static int input_fd;
@@ -155,10 +156,9 @@ static char *read_string(void)
static int read_proc_kallsyms(struct tep_handle *pevent)
{
unsigned int size;
+ char *buf;
size = read4(pevent);
- if (!size)
- return 0;
/*
* Just skip it, now that we configure libtraceevent to use the
* tools/perf/ symbol resolver.
@@ -170,11 +170,59 @@ static int read_proc_kallsyms(struct tep_handle *pevent)
* payload", so that older tools can continue reading it and interpret
* it as "no kallsyms payload is present".
*/
- lseek(input_fd, size, SEEK_CUR);
+ /* Write kallsyms section with empty payload if no data */
+ if (!size) {
+ if (trace_dat_fp && !trace_dat_write_failed) {
+ unsigned short section_id = to_file_u16(pevent, TRACE_DAT_SECTION_KALLSYMS);
+ unsigned short flags = to_file_u16(pevent, 0);
+ unsigned int string_id = to_file_u32(pevent, STRID_KALLSYMS);
+ unsigned long long section_size = to_file_u64(pevent, sizeof(unsigned int));
+ unsigned int kallsyms_data = to_file_u32(pevent, 0);
+
+ trace_dat_kallsyms_offset = ftell(trace_dat_fp);
+ if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp) ||
+ !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp) ||
+ !fwrite(&kallsyms_data, sizeof(unsigned int), 1, trace_dat_fp)) {
+ pr_warning("Failed to write trace.dat kallsyms section\n");
+ trace_dat_write_failed = true;
+ }
+ }
+ return 0;
+ }
+ buf = malloc(size);
+ if (buf == NULL)
+ return -1;
+ if (do_read(buf, size) < 0) {
+ free(buf);
+ return -1;
+ }
trace_data_size += size;
+ /* Write kallsyms section with data */
+ if (trace_dat_fp && !trace_dat_write_failed) {
+ unsigned short section_id = to_file_u16(pevent, TRACE_DAT_SECTION_KALLSYMS);
+ unsigned int string_id = to_file_u32(pevent, STRID_KALLSYMS);
+ unsigned long long section_size = to_file_u64(pevent, sizeof(unsigned int) + size);
+ unsigned short flags = to_file_u16(pevent, 0);
+ unsigned int size_out = to_file_u32(pevent, size);
+
+ trace_dat_kallsyms_offset = ftell(trace_dat_fp);
+ if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp) ||
+ !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp) ||
+ !fwrite(&size_out, sizeof(unsigned int), 1, trace_dat_fp) ||
+ !fwrite(buf, 1, size, trace_dat_fp)) {
+ pr_warning("Failed to write trace.dat kallsyms section\n");
+ trace_dat_write_failed = true;
+ }
+ }
+ free(buf);
return 0;
}
+
static int read_ftrace_printk(struct tep_handle *pevent)
{
unsigned int size;
@@ -210,6 +258,13 @@ static int read_ftrace_printk(struct tep_handle *pevent)
static int read_header_files(struct tep_handle *pevent)
{
unsigned long long size;
+ unsigned long long header_page_size;
+ unsigned long long header_event_size;
+ char *header_event;
+ unsigned short section_id;
+ unsigned short flags;
+ unsigned int string_id;
+ unsigned long long section_size;
char *header_page;
char buf[BUFSIZ];
ssize_t ret = 0;
@@ -224,6 +279,7 @@ static int read_header_files(struct tep_handle *pevent)
size = read8(pevent);
+ header_page_size = size;
header_page = malloc(size);
if (header_page == NULL)
return -1;
@@ -242,19 +298,63 @@ static int read_header_files(struct tep_handle *pevent)
*/
tep_set_long_size(pevent, tep_get_header_page_size(pevent));
}
- free(header_page);
- if (do_read(buf, 13) < 0)
+ if (do_read(buf, 13) < 0) {
+ free(header_page);
return -1;
+ }
if (memcmp(buf, "header_event", 13) != 0) {
pr_debug("did not read header event");
+ free(header_page);
return -1;
}
size = read8(pevent);
- skip(size);
+ if (trace_dat_fp && !trace_dat_write_failed) {
+ unsigned long long header_page_size_out;
+ unsigned long long header_event_size_out;
+
+ header_event_size = size;
+ header_event = malloc(size);
+ if (header_event == NULL) {
+ free(header_page);
+ return -1;
+ }
+ if (do_read(header_event, size) < 0) {
+ free(header_page);
+ free(header_event);
+ return -1;
+ }
+ /* Write header_page and header_event to trace.dat */
+ section_id = to_file_u16(pevent, TRACE_DAT_SECTION_HEADER);
+ flags = to_file_u16(pevent, 0);
+ string_id = to_file_u32(pevent, STRID_HEADERS);
+ section_size = to_file_u64(pevent, 12 + 8 + header_page_size +
+ 13 + 8 + header_event_size);
+ header_page_size_out = to_file_u64(pevent, header_page_size);
+ header_event_size_out = to_file_u64(pevent, header_event_size);
+
+ trace_dat_header_info_offset = ftell(trace_dat_fp);
+ if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp) ||
+ !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp) ||
+ !fwrite("header_page\0", 1, 12, trace_dat_fp) ||
+ !fwrite(&header_page_size_out, sizeof(unsigned long long), 1, trace_dat_fp) ||
+ !fwrite(header_page, 1, header_page_size, trace_dat_fp) ||
+ !fwrite("header_event\0", 1, 13, trace_dat_fp) ||
+ !fwrite(&header_event_size_out, sizeof(unsigned long long), 1, trace_dat_fp) ||
+ !fwrite(header_event, 1, header_event_size, trace_dat_fp)) {
+ pr_warning("Failed to write trace.dat header section\n");
+ trace_dat_write_failed = true;
+ }
+ free(header_event);
+ } else {
+ skip(size);
+ }
+ free(header_page);
return ret;
}
@@ -274,6 +374,15 @@ static int read_ftrace_file(struct tep_handle *pevent, unsigned long long size)
pr_debug("error reading ftrace file.\n");
goto out;
}
+ if (trace_dat_fp && !trace_dat_write_failed) {
+ unsigned long long size_out = to_file_u64(pevent, size);
+
+ if (!fwrite(&size_out, sizeof(unsigned long long), 1, trace_dat_fp) ||
+ !fwrite(buf, 1, size, trace_dat_fp)) {
+ pr_warning("Failed to write trace.dat ftrace event formats section\n");
+ trace_dat_write_failed = true;
+ }
+ }
ret = parse_ftrace_file(pevent, buf, size);
if (ret < 0)
@@ -298,6 +407,15 @@ static int read_event_file(struct tep_handle *pevent, char *sys,
ret = do_read(buf, size);
if (ret < 0)
goto out;
+ if (trace_dat_fp && !trace_dat_write_failed) {
+ unsigned long long size_out = to_file_u64(pevent, size);
+
+ if (!fwrite(&size_out, sizeof(unsigned long long), 1, trace_dat_fp) ||
+ !fwrite(buf, 1, size, trace_dat_fp)) {
+ pr_warning("Failed to write trace.dat event formats section\n");
+ trace_dat_write_failed = true;
+ }
+ }
ret = parse_event_file(pevent, buf, size, sys);
if (ret < 0)
@@ -313,8 +431,39 @@ static int read_ftrace_files(struct tep_handle *pevent)
int count;
int i;
int ret;
+ long section_size_pos = 0;
+ long count_pos = 0;
+ unsigned long long section_size = 0;
+ long end_pos;
count = read4(pevent);
+ /* Write ftrace formats section to trace.dat output file */
+ if (trace_dat_fp && !trace_dat_write_failed) {
+ unsigned short section_id = to_file_u16(pevent, TRACE_DAT_SECTION_FTRACE);
+ unsigned short flags = to_file_u16(pevent, 0);
+ unsigned int string_id = to_file_u32(pevent, STRID_FTRACE_FORMATS);
+ unsigned int count_out = to_file_u32(pevent, count);
+
+ section_size = to_file_u64(pevent, 0);
+ trace_dat_ftrace_format_offset = ftell(trace_dat_fp);
+
+ if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp)) {
+ pr_warning("Failed to write trace.dat ftrace event formats section\n");
+ trace_dat_write_failed = true;
+ }
+ section_size_pos = ftell(trace_dat_fp);
+ if (!fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp)) {
+ pr_warning("Failed to write trace.dat ftrace event formats section\n");
+ trace_dat_write_failed = true;
+ }
+ count_pos = ftell(trace_dat_fp);
+ if (!fwrite(&count_out, sizeof(unsigned int), 1, trace_dat_fp)) {
+ pr_warning("Failed to write trace.dat ftrace event formats section\n");
+ trace_dat_write_failed = true;
+ }
+ }
for (i = 0; i < count; i++) {
size = read8(pevent);
@@ -322,6 +471,18 @@ static int read_ftrace_files(struct tep_handle *pevent)
if (ret)
return ret;
}
+ /* Fill in section size after writing all ftrace files */
+ if (trace_dat_fp && !trace_dat_write_failed) {
+ end_pos = ftell(trace_dat_fp);
+ section_size = to_file_u64(pevent, end_pos - count_pos);
+ if (fseek(trace_dat_fp, section_size_pos, SEEK_SET) < 0 ||
+ !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp) ||
+ fseek(trace_dat_fp, end_pos, SEEK_SET) < 0) {
+ pr_warning("Failed to write trace.dat ftrace event formats section\n");
+ trace_dat_write_failed = true;
+ }
+ }
+
return 0;
}
@@ -333,8 +494,38 @@ static int read_event_files(struct tep_handle *pevent)
int count;
int i,x;
ssize_t ret;
+ long section_size_pos = 0;
+ long sys_count_pos = 0;
+ unsigned long long section_size = 0;
+ long end_pos;
systems = read4(pevent);
+ /* Write event formats section to trace.dat output file */
+ if (trace_dat_fp && !trace_dat_write_failed) {
+ unsigned short section_id = to_file_u16(pevent, TRACE_DAT_SECTION_EVENTS);
+ unsigned short flags = to_file_u16(pevent, 0);
+ unsigned int string_id = to_file_u32(pevent, STRID_EVENT_FORMATS);
+ unsigned int systems_out = to_file_u32(pevent, systems);
+
+ section_size = to_file_u64(pevent, 0);
+ trace_dat_events_format_offset = ftell(trace_dat_fp);
+ if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp)) {
+ pr_warning("Failed to write trace.dat event formats section\n");
+ trace_dat_write_failed = true;
+ }
+ section_size_pos = ftell(trace_dat_fp);
+ if (!fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp)) {
+ pr_warning("Failed to write trace.dat event formats section\n");
+ trace_dat_write_failed = true;
+ }
+ sys_count_pos = ftell(trace_dat_fp);
+ if (!fwrite(&systems_out, sizeof(unsigned int), 1, trace_dat_fp)) {
+ pr_warning("Failed to write trace.dat event formats section\n");
+ trace_dat_write_failed = true;
+ }
+ }
for (i = 0; i < systems; i++) {
sys = read_string();
@@ -342,6 +533,15 @@ static int read_event_files(struct tep_handle *pevent)
return -1;
count = read4(pevent);
+ if (trace_dat_fp && !trace_dat_write_failed) {
+ unsigned int count_out = to_file_u32(pevent, count);
+
+ if (!fwrite(sys, 1, strlen(sys) + 1, trace_dat_fp) ||
+ !fwrite(&count_out, sizeof(unsigned int), 1, trace_dat_fp)) {
+ pr_warning("Failed to write trace.dat event formats section\n");
+ trace_dat_write_failed = true;
+ }
+ }
for (x=0; x < count; x++) {
size = read8(pevent);
@@ -353,6 +553,18 @@ static int read_event_files(struct tep_handle *pevent)
}
free(sys);
}
+ /* Fill in section size after writing all event files */
+ if (trace_dat_fp && !trace_dat_write_failed) {
+ end_pos = ftell(trace_dat_fp);
+ section_size = to_file_u64(pevent, end_pos - sys_count_pos);
+ if (fseek(trace_dat_fp, section_size_pos, SEEK_SET) < 0 ||
+ !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp) ||
+ fseek(trace_dat_fp, end_pos, SEEK_SET) < 0) {
+ pr_warning("Failed to write trace.dat event formats section\n");
+ trace_dat_write_failed = true;
+ }
+ }
+
return 0;
}
@@ -364,8 +576,28 @@ static int read_saved_cmdline(struct tep_handle *pevent)
/* it can have 0 size */
size = read8(pevent);
- if (!size)
+ /* Write cmdlines section with empty payload if no data */
+ if (!size) {
+ if (trace_dat_fp && !trace_dat_write_failed) {
+ unsigned short section_id = to_file_u16(pevent, TRACE_DAT_SECTION_CMDLINE);
+ unsigned short flags = to_file_u16(pevent, 0);
+ unsigned int string_id = to_file_u32(pevent, STRID_CMDLINES);
+ unsigned long long section_size =
+ to_file_u64(pevent, sizeof(unsigned long long));
+ unsigned long long section_data = to_file_u64(pevent, 0);
+
+ trace_dat_cmdline_offset = ftell(trace_dat_fp);
+ if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp) ||
+ !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp) ||
+ !fwrite(§ion_data, sizeof(unsigned long long), 1, trace_dat_fp)) {
+ pr_warning("Failed to write trace.dat cmdlines section\n");
+ trace_dat_write_failed = true;
+ }
+ }
return 0;
+ }
if (size == ULLONG_MAX) {
pr_debug("invalid saved cmdline size");
@@ -383,6 +615,27 @@ static int read_saved_cmdline(struct tep_handle *pevent)
pr_debug("error reading saved cmdlines\n");
goto out;
}
+ /* Write cmdlines section with data */
+ if (trace_dat_fp && !trace_dat_write_failed) {
+ unsigned short section_id = to_file_u16(pevent, TRACE_DAT_SECTION_CMDLINE);
+ unsigned short flags = to_file_u16(pevent, 0);
+ unsigned int string_id = to_file_u32(pevent, STRID_CMDLINES);
+ unsigned long long section_size =
+ to_file_u64(pevent, sizeof(unsigned long long) + size);
+ unsigned long long size_out = to_file_u64(pevent, size);
+
+ trace_dat_cmdline_offset = ftell(trace_dat_fp);
+ if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) ||
+ !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp) ||
+ !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp) ||
+ !fwrite(&size_out, sizeof(unsigned long long), 1, trace_dat_fp) ||
+ !fwrite(buf, 1, size, trace_dat_fp)) {
+ pr_warning("Failed to write trace.dat cmdlines section\n");
+ trace_dat_write_failed = true;
+ }
+ }
+
buf[ret] = '\0';
parse_saved_cmdline(pevent, buf, size);
@@ -407,6 +660,7 @@ ssize_t trace_report(int fd, struct trace_event *tevent, bool __repipe)
int file_page_size;
struct tep_handle *pevent = NULL;
int err;
+ char magic_buf[10];
repipe = __repipe;
input_fd = fd;
@@ -418,12 +672,17 @@ ssize_t trace_report(int fd, struct trace_event *tevent, bool __repipe)
return -1;
}
+ if (trace_dat_fp)
+ memcpy(magic_buf, buf, 3);
+
if (do_read(buf, 7) < 0)
return -1;
if (memcmp(buf, "tracing", 7) != 0) {
pr_debug("not a trace file (missing 'tracing' tag)");
return -1;
}
+ if (trace_dat_fp)
+ memcpy(magic_buf + 3, buf, 7);
version = read_string();
if (version == NULL)
@@ -460,6 +719,34 @@ ssize_t trace_report(int fd, struct trace_event *tevent, bool __repipe)
tep_set_long_size(pevent, file_long_size);
tep_set_page_size(pevent, file_page_size);
+ /* Write initial file header to trace.dat */
+ if (trace_dat_fp && !trace_dat_write_failed) {
+ unsigned char endian = file_bigendian;
+ unsigned char long_size = file_long_size;
+ unsigned int page_size = to_file_u32(pevent, file_page_size);
+ unsigned long long placeholder = to_file_u64(pevent, 0);
+
+ if (!fwrite(magic_buf, 1, 10, trace_dat_fp) || /* magic + "tracing" */
+ !fwrite(TRACE_DAT_VERSION, 1, 2, trace_dat_fp) ||
+ !fwrite(&endian, 1, 1, trace_dat_fp) ||
+ !fwrite(&long_size, 1, 1, trace_dat_fp) ||
+ !fwrite(&page_size, sizeof(unsigned int), 1, trace_dat_fp) ||
+ !fwrite("none", 1, 4, trace_dat_fp) ||
+ !fwrite("\0", 1, 1, trace_dat_fp) ||
+ !fwrite("\0", 1, 1, trace_dat_fp)) {
+ pr_warning("Failed to write trace.dat initalfile header\n");
+ trace_dat_write_failed = true;
+ }
+
+ if (!trace_dat_write_failed) {
+ trace_dat_options_offset = ftell(trace_dat_fp);
+ if (!fwrite(&placeholder, sizeof(unsigned long long), 1, trace_dat_fp)) {
+ pr_warning("Failed to write trace.dat initial file header\n");
+ trace_dat_write_failed = true;
+ }
+ }
+ }
+
err = read_header_files(pevent);
if (err)
goto out;
@@ -480,6 +767,12 @@ ssize_t trace_report(int fd, struct trace_event *tevent, bool __repipe)
if (err)
goto out;
}
+ /* Write strings section to trace.dat output file */
+ if (trace_dat_fp && !trace_dat_write_failed) {
+ err = trace_dat__write_strings_section(pevent);
+ if (err)
+ goto out;
+ }
size = trace_data_size;
repipe = false;
--
2.47.1
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [RFC PATCH v4 3/5] perf data-convert: Add perf.data to trace.dat conversion backend
2026-08-22 6:27 [RFC PATCH v4 0/5] Add perf.data tracepoint events to trace.dat conversion Tanushree Shah
2026-08-22 6:27 ` [RFC PATCH v4 1/5] perf/trace-dat: Add trace.dat export infrastructure Tanushree Shah
2026-08-22 6:27 ` [RFC PATCH v4 2/5] perf/trace-event: Write trace.dat metadata sections during parsing Tanushree Shah
@ 2026-08-22 6:27 ` Tanushree Shah
2026-08-22 6:38 ` sashiko-bot
2026-08-22 6:27 ` [RFC PATCH v4 4/5] perf data: Add --to-trace-dat option for converting perf.data tracepoint events into trace.dat format Tanushree Shah
2026-08-22 6:27 ` [RFC PATCH v4 5/5] perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat' Tanushree Shah
4 siblings, 1 reply; 11+ messages in thread
From: Tanushree Shah @ 2026-08-22 6:27 UTC (permalink / raw)
To: acme, jolsa, adrian.hunter, vmolnaro, mpetlan, tmricht, maddy,
irogers, namhyung
Cc: linux-perf-users, linuxppc-dev, atrajeev, hbathini, Tejas.Manhas1,
Tanushree.Shah, rostedt, Tanushree Shah
Add data-convert-trace.c implementing trace_convert__perf2dat() to
convert perf.data tracepoint events to trace.dat format.
process_sample_event() is invoked for each PERF_TYPE_TRACEPOINT sample
during perf_session__process_events(), storing raw event bytes per-cpu
via trace_dat__collect_cpu_event().
Once all samples are collected:
- trace_dat__write_options_section1() writes the OPTIONS section with
CPUCOUNT, TRACECLOCK, HEADER_INFO, FTRACE_EVENTS, EVENT_FORMATS,
KALLSYMS, CMDLINES and DONE options.
- trace_dat__write_options_section2() writes the OPTIONS section with
BUFFER option holding per-cpu data offset placeholders and the DONE
option.
- trace_dat__write_flyrecord_section() builds ring buffer pages
per-cpu and patches BUFFER option with final offsets and sizes,
Per-cpu buffers are sized to tep_get_page_size() from the session
tep handle and released on all exit paths.
Register .attr, .feature, and .tracing_data callbacks on the tool,
delegating to the standard perf_event__process_attr(),
perf_event__process_feature(), and perf_event__process_tracing_data()
helpers, so these records are processed as they arrive in the stream
instead of being left as default no-op stubs.
Defer per-CPU buffer initialisation to the first call of the .sample
callback instead of doing it eagerly after session open. By this
point, pipe-mode ordering guarantees (record__synthesize()) that
attr, feature, and tracing_data records have already been written and
processed, so nr_cpus_online and the tep page size are guaranteed to
be populated from the recording machine's own stream data. This is
required because the converting host may differ in architecture, CPU
count, or page size from the machine that recorded the data, so these
values must never be taken from local sysconf() or similar.
The output path honours --force: without it, the destination file is
opened with O_CREAT|O_EXCL so an existing file is detected atomically
and conversion fails with -EEXIST rather than silently overwriting
prior output.
Signed-off-by: Tanushree Shah <tshah@linux.ibm.com>
---
tools/perf/util/Build | 1 +
tools/perf/util/data-convert-trace.c | 243 +++++++++++++++++++++++++++
tools/perf/util/data-convert.h | 4 +
3 files changed, 248 insertions(+)
create mode 100644 tools/perf/util/data-convert-trace.c
diff --git a/tools/perf/util/Build b/tools/perf/util/Build
index 12f5bbe32770..40db8e3c01c4 100644
--- a/tools/perf/util/Build
+++ b/tools/perf/util/Build
@@ -237,6 +237,7 @@ ifeq ($(CONFIG_LIBTRACEEVENT),y)
endif
perf-util-y += data-convert-json.o
+perf-util-$(CONFIG_LIBTRACEEVENT) += data-convert-trace.o
perf-util-y += scripting-engines/
diff --git a/tools/perf/util/data-convert-trace.c b/tools/perf/util/data-convert-trace.c
new file mode 100644
index 000000000000..445479fae888
--- /dev/null
+++ b/tools/perf/util/data-convert-trace.c
@@ -0,0 +1,243 @@
+// SPDX-License-Identifier: GPL-2.0-or-later
+/*
+ * Copyright 2026, IBM Corporation
+ * Author: Tanushree Shah <tshah@linux.ibm.com>
+ *
+ * data-convert-trace.c
+ *
+ * Implements perf.data to trace.dat format conversion for tracepoint events.
+ */
+
+#include <errno.h>
+#include <inttypes.h>
+#include <fcntl.h>
+#include <stdio.h>
+#include <string.h>
+#include <unistd.h>
+#include <linux/compiler.h>
+#include <linux/err.h>
+
+#include "data-convert.h"
+#include "session.h"
+#include "evsel.h"
+#include "tool.h"
+#include "debug.h"
+#include "trace-dat.h"
+#include "trace-event.h"
+#include "event.h"
+#include "sample.h"
+#include "evlist.h"
+
+struct trace_convert {
+ struct perf_tool tool;
+ u64 events_count;
+};
+
+/* Session handle and init flag used for lazy CPU buffer init in pipe mode */
+static struct perf_session *trace_dat_session;
+static bool cpu_buffers_initialized;
+
+/* Store raw tracepoint event data in per-cpu buffer for trace.dat flyrecord */
+static int process_sample_event(const struct perf_tool *tool,
+ union perf_event *event __maybe_unused,
+ struct perf_sample *sample,
+ struct machine *machine __maybe_unused)
+{
+ struct trace_convert *tc = container_of(tool, struct trace_convert, tool);
+ struct evsel *evsel = sample->evsel;
+
+ /* Only process tracepoint events */
+ if (!trace_dat_fp || sample->raw_size == 0 ||
+ !evsel ||
+ evsel->core.attr.type != PERF_TYPE_TRACEPOINT)
+ return 0;
+
+ /*
+ * In pipe mode, CPU count and page size arrive via feature/tracing_data
+ * records before the first sample; initialize buffers lazily on first sample.
+ */
+ if (!cpu_buffers_initialized) {
+ int nr_cpus = trace_dat_session->header.env.nr_cpus_online;
+
+ if (trace_dat_session->tevent.pevent)
+ trace_dat_page_size = tep_get_page_size(trace_dat_session->tevent.pevent);
+
+ /*
+ * nr_cpus and page size MUST come from the recording
+ * machine's own stream data (feature/tracing_data records
+ * never from local sysconf() - the converting host may be
+ * a different arch/cpu-count/page-size than the machine
+ * that recorded the data. Pipe-mode ordering guarantees
+ * (record__synthesize()) that attr -> feature ->
+ * tracing_data are always written before the first sample,
+ * so by the time we're in this callback, both values
+ * should already be populated by process_feature()/
+ * process_tracing_data().
+ */
+ if (nr_cpus <= 0 || trace_dat_page_size == 0) {
+ pr_err("Malformed pipe stream: missing feature/tracing_data before first sample\n");
+ return -EINVAL;
+ }
+
+ if (trace_dat__init_cpu_buffers(nr_cpus) < 0) {
+ pr_err("Failed to initialize CPU buffers\n");
+ return -ENOMEM;
+ }
+ cpu_buffers_initialized = true;
+ }
+
+ if (trace_dat__collect_cpu_event(sample->cpu, sample->time,
+ sample->raw_data, sample->raw_size) < 0) {
+ pr_err("Failed to collect CPU event\n");
+ return -ENOMEM;
+ }
+ tc->events_count++;
+
+ return 0;
+}
+
+/* Process event attributes for pipe mode */
+static int process_attr(const struct perf_tool *tool __maybe_unused,
+ union perf_event *event,
+ struct evlist **pevlist)
+{
+ return perf_event__process_attr(tool, event, pevlist);
+}
+
+/* Process feature records for pipe mode */
+static int process_feature(const struct perf_tool *tool __maybe_unused,
+ struct perf_session *session,
+ union perf_event *event)
+{
+ return perf_event__process_feature(tool, session, event);
+}
+
+/* Process tracing data for pipe mode */
+static int process_tracing_data(const struct perf_tool *tool __maybe_unused,
+ struct perf_session *session,
+ union perf_event *event)
+{
+ return perf_event__process_tracing_data(tool, session, event);
+}
+
+/* Convert perf.data tracepoint events to trace.dat format */
+int trace_convert__perf2dat(const char *input, const char *to_trace,
+ struct perf_data_convert_opts *opts)
+{
+ struct perf_session *session;
+ struct trace_convert tc = {
+ .events_count = 0,
+ };
+ struct perf_data data = {
+ .path = input,
+ .mode = PERF_DATA_MODE_READ,
+ .force = opts->force,
+ };
+ int ret = -EINVAL;
+
+ cpu_buffers_initialized = false;
+
+ /* Initialize tool with all required callbacks */
+ perf_tool__init(&tc.tool, /*ordered_events=*/true);
+ tc.tool.sample = process_sample_event;
+ tc.tool.attr = process_attr;
+ tc.tool.feature = process_feature;
+ tc.tool.tracing_data = process_tracing_data;
+
+ if (!opts->force) {
+ int fd = open(to_trace, O_WRONLY | O_CREAT | O_EXCL, 0644);
+
+ if (fd < 0) {
+ if (errno == EEXIST)
+ pr_err("Output file '%s' already exists. Use --force to overwrite.\n",
+ to_trace);
+ else
+ pr_err("Failed to open output file '%s': %s\n",
+ to_trace, strerror(errno));
+ return -errno;
+ }
+ trace_dat_fp = fdopen(fd, "wb");
+ if (!trace_dat_fp) {
+ int err = errno;
+
+ close(fd);
+ pr_err("Failed to open output file '%s': %s\n",
+ to_trace, strerror(err));
+ ret = -err;
+ goto out_close;
+ }
+ } else {
+ trace_dat_fp = fopen(to_trace, "wb");
+ if (!trace_dat_fp) {
+ pr_err("Failed to open output file: %s\n", to_trace);
+ return -EINVAL;
+ }
+ }
+
+ /* Open perf.data session - this writes trace.dat metadata sections */
+ session = perf_session__new(&data, &tc.tool);
+ if (IS_ERR(session)) {
+ pr_err("Failed to open perf.data file\n");
+ ret = PTR_ERR(session);
+ goto out_close;
+ }
+
+ /* Stash session for lazy CPU buffer init on first sample (pipe and normal mode) */
+ trace_dat_session = session;
+
+ /* Process all events - collects raw data per-cpu */
+ ret = perf_session__process_events(session);
+ if (ret < 0) {
+ pr_err("Failed to process events\n");
+ goto out_delete;
+ }
+
+ /* Skip file creation if no tracepoint events found */
+ if (tc.events_count == 0) {
+ pr_warning("No tracepoint events found in '%s', skipping trace.dat creation\n",
+ input);
+ ret = -EINVAL;
+ goto out_delete;
+ }
+
+ /* Write trace.dat options and flyrecord sections */
+ if (trace_dat__write_options_section1(session->tevent.pevent) < 0 ||
+ trace_dat_write_failed) {
+ pr_err("Failed to write options section1\n");
+ ret = -EIO;
+ goto out_delete;
+ }
+ if (trace_dat__write_options_section2(session->tevent.pevent) < 0 ||
+ trace_dat_write_failed) {
+ pr_err("Failed to write options section2\n");
+ ret = -EIO;
+ goto out_delete;
+ }
+ if (trace_dat__write_flyrecord_section(session->tevent.pevent) < 0 ||
+ trace_dat_write_failed) {
+ pr_err("Failed to write flyrecord section\n");
+ ret = -EIO;
+ goto out_delete;
+ }
+
+ pr_info("[ perf data convert: Converted '%s' into trace.dat format '%s' ]\n",
+ input, to_trace);
+ pr_info("[ perf data convert: Converted %llu events ]\n",
+ (unsigned long long)tc.events_count);
+
+ ret = 0;
+
+out_delete:
+ if (cpu_buffers_initialized)
+ trace_dat__free_cpu_buffers();
+ perf_session__delete(session);
+ trace_dat_session = NULL;
+out_close:
+ if (trace_dat_fp) {
+ fclose(trace_dat_fp);
+ trace_dat_fp = NULL;
+ }
+ if (ret != 0)
+ unlink(to_trace);
+ return ret;
+}
diff --git a/tools/perf/util/data-convert.h b/tools/perf/util/data-convert.h
index a96240f15671..f041c2325226 100644
--- a/tools/perf/util/data-convert.h
+++ b/tools/perf/util/data-convert.h
@@ -19,4 +19,8 @@ int bt_convert__perf2ctf(const char *input_name, const char *to_ctf,
int bt_convert__perf2json(const char *input_name, const char *to_ctf,
struct perf_data_convert_opts *opts);
+#ifdef HAVE_LIBTRACEEVENT
+int trace_convert__perf2dat(const char *input, const char *to_trace,
+ struct perf_data_convert_opts *opts);
+#endif /* HAVE_LIBTRACEEVENT */
#endif /* __DATA_CONVERT_H */
--
2.47.1
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [RFC PATCH v4 4/5] perf data: Add --to-trace-dat option for converting perf.data tracepoint events into trace.dat format
2026-08-22 6:27 [RFC PATCH v4 0/5] Add perf.data tracepoint events to trace.dat conversion Tanushree Shah
` (2 preceding siblings ...)
2026-08-22 6:27 ` [RFC PATCH v4 3/5] perf data-convert: Add perf.data to trace.dat conversion backend Tanushree Shah
@ 2026-08-22 6:27 ` Tanushree Shah
2026-08-22 6:44 ` sashiko-bot
2026-08-22 6:27 ` [RFC PATCH v4 5/5] perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat' Tanushree Shah
4 siblings, 1 reply; 11+ messages in thread
From: Tanushree Shah @ 2026-08-22 6:27 UTC (permalink / raw)
To: acme, jolsa, adrian.hunter, vmolnaro, mpetlan, tmricht, maddy,
irogers, namhyung
Cc: linux-perf-users, linuxppc-dev, atrajeev, hbathini, Tejas.Manhas1,
Tanushree.Shah, rostedt, Tanushree Shah
Add new command-line option to perf data convert for generating
trace.dat output files.
The --to-trace-dat option:
- Accepts output filename for trace.dat format
- Mutually exclusive with --to-ctf and --to-json
- Calls trace_convert__perf2dat() to perform conversion
Usage:
$ perf record -e sched:* -a sleep 1
$ perf data convert --to-trace-dat=trace.dat
$ trace-cmd report trace.dat
Document the new option in Documentation/perf-data.txt alongside
the existing --to-ctf and --to-json entries.
Honor the --time filtering option in trace_convert__perf2dat() the
same way the JSON and CTF converters do, using
perf_time__parse_for_ranges() and perf_time__ranges_skip_sample()
to skip samples outside the requested time range.
Signed-off-by: Tanushree Shah <tshah@linux.ibm.com>
---
tools/perf/Documentation/perf-data.txt | 7 +++++
tools/perf/builtin-data.c | 43 ++++++++++++++++++++++++++
tools/perf/util/data-convert-trace.c | 21 +++++++++++++
tools/perf/util/trace-dat.c | 20 +++++++++---
tools/perf/util/trace-dat.h | 1 +
5 files changed, 87 insertions(+), 5 deletions(-)
diff --git a/tools/perf/Documentation/perf-data.txt b/tools/perf/Documentation/perf-data.txt
index 20f178d61ed7..578ba6357183 100644
--- a/tools/perf/Documentation/perf-data.txt
+++ b/tools/perf/Documentation/perf-data.txt
@@ -30,6 +30,13 @@ OPTIONS for 'convert'
--to-json::
Triggers JSON conversion. Specify the JSON filename to output.
+--to-trace-dat::
+ Triggers trace.dat conversion. Converts perf.data tracepoint events
+ to trace.dat format v7, compatible with trace-cmd and KernelShark.
+ Only PERF_TYPE_TRACEPOINT events are converted. Specify the
+ trace.dat filename to output. Requires libtraceevent support.
+ Mutually exclusive with --to-ctf and --to-json.
+
--tod::
Convert time to wall clock time.
diff --git a/tools/perf/builtin-data.c b/tools/perf/builtin-data.c
index 1dd73ed6bdcb..c9c863197d02 100644
--- a/tools/perf/builtin-data.c
+++ b/tools/perf/builtin-data.c
@@ -30,6 +30,11 @@ static const char *data_usage[] = {
static const char *to_json;
static const char *to_ctf;
+
+#ifdef HAVE_LIBTRACEEVENT
+static const char *trace_dat_output;
+#endif
+
static struct perf_data_convert_opts opts = {
.force = false,
.all = false,
@@ -46,6 +51,10 @@ static const struct option data_options[] = {
OPT_BOOLEAN(0, "all", &opts.all, "Convert all events"),
OPT_STRING(0, "time", &opts.time_str, "str",
"Time span of interest (start,stop)"),
+#ifdef HAVE_LIBTRACEEVENT
+ OPT_STRING(0, "to-trace-dat", &trace_dat_output,
+ "file", "Convert to trace.dat format using perf.data tracepoints"),
+#endif
OPT_END()
};
@@ -63,10 +72,44 @@ static int cmd_data_convert(int argc, const char **argv)
pr_err("You cannot specify both --to-ctf and --to-json.\n");
return -1;
}
+#ifdef HAVE_LIBTRACEEVENT
+ if (trace_dat_output && (to_json || to_ctf)) {
+ pr_err("You cannot specify --to-trace-dat with --to-ctf or --to-json.\n");
+ return -1;
+ }
+#endif
+
+#ifdef HAVE_LIBBABELTRACE_SUPPORT
+ #ifdef HAVE_LIBTRACEEVENT
+ if (!to_json && !to_ctf && !trace_dat_output) {
+ pr_err("You must specify one of --to-ctf, --to-json, or --to-trace-dat.\n");
+ return -1;
+ }
+ #else
if (!to_json && !to_ctf) {
pr_err("You must specify one of --to-ctf or --to-json.\n");
return -1;
}
+ #endif
+#else
+ #ifdef HAVE_LIBTRACEEVENT
+ if (!to_json && !trace_dat_output) {
+ pr_err("You must specify --to-json or --to-trace-dat.\n");
+ return -1;
+ }
+ #else
+ if (!to_json) {
+ pr_err("You must specify --to-json.\n");
+ return -1;
+ }
+ #endif
+#endif
+
+#ifdef HAVE_LIBTRACEEVENT
+ if (trace_dat_output)
+ return trace_convert__perf2dat(input_name ? input_name : "perf.data",
+ trace_dat_output, &opts);
+#endif
if (to_json)
return bt_convert__perf2json(input_name, to_json, &opts);
diff --git a/tools/perf/util/data-convert-trace.c b/tools/perf/util/data-convert-trace.c
index 445479fae888..425b3aaa3f02 100644
--- a/tools/perf/util/data-convert-trace.c
+++ b/tools/perf/util/data-convert-trace.c
@@ -22,6 +22,7 @@
#include "evsel.h"
#include "tool.h"
#include "debug.h"
+#include "time-utils.h"
#include "trace-dat.h"
#include "trace-event.h"
#include "event.h"
@@ -31,6 +32,10 @@
struct trace_convert {
struct perf_tool tool;
u64 events_count;
+ struct perf_time_interval *ptime_range;
+ int range_size;
+ int range_num;
+ u64 skipped;
};
/* Session handle and init flag used for lazy CPU buffer init in pipe mode */
@@ -86,6 +91,11 @@ static int process_sample_event(const struct perf_tool *tool,
cpu_buffers_initialized = true;
}
+ if (perf_time__ranges_skip_sample(tc->ptime_range, tc->range_num, sample->time)) {
+ tc->skipped++;
+ return 0;
+ }
+
if (trace_dat__collect_cpu_event(sample->cpu, sample->time,
sample->raw_data, sample->raw_size) < 0) {
pr_err("Failed to collect CPU event\n");
@@ -185,6 +195,15 @@ int trace_convert__perf2dat(const char *input, const char *to_trace,
/* Stash session for lazy CPU buffer init on first sample (pipe and normal mode) */
trace_dat_session = session;
+ if (opts->time_str) {
+ ret = perf_time__parse_for_ranges(opts->time_str, session,
+ &tc.ptime_range,
+ &tc.range_size,
+ &tc.range_num);
+ if (ret < 0)
+ goto out_delete;
+ }
+
/* Process all events - collects raw data per-cpu */
ret = perf_session__process_events(session);
if (ret < 0) {
@@ -230,6 +249,8 @@ int trace_convert__perf2dat(const char *input, const char *to_trace,
out_delete:
if (cpu_buffers_initialized)
trace_dat__free_cpu_buffers();
+ if (tc.ptime_range)
+ zfree(&tc.ptime_range);
perf_session__delete(session);
trace_dat_session = NULL;
out_close:
diff --git a/tools/perf/util/trace-dat.c b/tools/perf/util/trace-dat.c
index f71e03716e27..1b236745bcac 100644
--- a/tools/perf/util/trace-dat.c
+++ b/tools/perf/util/trace-dat.c
@@ -194,6 +194,7 @@ static int trace_dat__write_cpu_dat(FILE *fp, struct tep_handle *pevent,
int page_size_used = 0;
int ret = 0;
int i, j;
+ unsigned long long page_base_ts;
file_offset = ftell(fp);
*file_offset_out = file_offset;
@@ -212,6 +213,7 @@ static int trace_dat__write_cpu_dat(FILE *fp, struct tep_handle *pevent,
}
base_ts = cpu_events->events[0].ts;
+ page_base_ts = base_ts;
for (i = 0; i < cpu_events->count; i++) {
struct cpu_event *event = &cpu_events->events[i];
@@ -234,8 +236,10 @@ static int trace_dat__write_cpu_dat(FILE *fp, struct tep_handle *pevent,
extend_size = TRACE_DAT_RECORD_TIME_EXTEND_SIZE;
extend = calloc(1, extend_size);
- if (!extend)
- return -ENOMEM;
+ if (!extend) {
+ ret = -ENOMEM;
+ goto out_free;
+ }
if (tep_is_file_bigendian(pevent)) {
extend_hdr = (time_delta & TRACE_DAT_RECORD_TIME_MASK) |
@@ -245,7 +249,7 @@ static int trace_dat__write_cpu_dat(FILE *fp, struct tep_handle *pevent,
extend_hdr = ((time_delta & TRACE_DAT_RECORD_TIME_MASK) <<
TRACE_DAT_RECORD_TIME_SHIFT) |
TRACE_DAT_RECORD_TYPE_TIME_EXTEND;
- delta_upper = time_delta >> TRACE_DAT_RECORD_TIME_SHIFT;
+ delta_upper = time_delta >> 27;
}
extend_hdr = to_file_u32(pevent, extend_hdr); /* still needed */
delta_upper = to_file_u32(pevent, delta_upper); /* still needed */
@@ -288,7 +292,7 @@ static int trace_dat__write_cpu_dat(FILE *fp, struct tep_handle *pevent,
/* Check page fit BEFORE allocating data record */
if (page_size_used + needed_size >
trace_dat_page_size - TRACE_DAT_RECORD_HEADER_SIZE) {
- ret = trace_dat__write_page(fp, pevent, base_ts,
+ ret = trace_dat__write_page(fp, pevent, page_base_ts,
page_records, page_rec_sizes,
nr_page_recs);
@@ -298,6 +302,7 @@ static int trace_dat__write_cpu_dat(FILE *fp, struct tep_handle *pevent,
nr_page_recs = 0;
page_size_used = 0;
base_ts = event->ts;
+ page_base_ts = event->ts;
if (ret < 0) {
free(extend);
@@ -312,6 +317,7 @@ static int trace_dat__write_cpu_dat(FILE *fp, struct tep_handle *pevent,
extend = NULL;
extend_size = 0;
time_delta = 0;
+ base_ts = event->ts;
}
if (tep_is_file_bigendian(pevent))
@@ -387,10 +393,11 @@ static int trace_dat__write_cpu_dat(FILE *fp, struct tep_handle *pevent,
page_rec_sizes[nr_page_recs] = data_rec_size;
nr_page_recs++;
page_size_used += data_rec_size;
+ base_ts = event->ts;
}
if (nr_page_recs > 0) {
- ret = trace_dat__write_page(fp, pevent, base_ts,
+ ret = trace_dat__write_page(fp, pevent, page_base_ts,
page_records, page_rec_sizes, nr_page_recs);
}
out_free:
@@ -861,6 +868,9 @@ void trace_dat__free_cpu_buffers(void)
for (cpu = 0; cpu < trace_dat_nr_cpus; cpu++) {
int i;
+ if (!trace_cpu_data[cpu].events)
+ continue;
+
for (i = 0; i < trace_cpu_data[cpu].count; i++)
free(trace_cpu_data[cpu].events[i].raw);
free(trace_cpu_data[cpu].events);
diff --git a/tools/perf/util/trace-dat.h b/tools/perf/util/trace-dat.h
index 5985083b275a..40041dac60d4 100644
--- a/tools/perf/util/trace-dat.h
+++ b/tools/perf/util/trace-dat.h
@@ -11,6 +11,7 @@
#include <stdbool.h>
#include <event-parse.h>
#include <byteswap.h>
+#include <stdint.h>
#include "util.h"
/* trace.dat file format version */
--
2.47.1
^ permalink raw reply related [flat|nested] 11+ messages in thread
* [RFC PATCH v4 5/5] perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat'
2026-08-22 6:27 [RFC PATCH v4 0/5] Add perf.data tracepoint events to trace.dat conversion Tanushree Shah
` (3 preceding siblings ...)
2026-08-22 6:27 ` [RFC PATCH v4 4/5] perf data: Add --to-trace-dat option for converting perf.data tracepoint events into trace.dat format Tanushree Shah
@ 2026-08-22 6:27 ` Tanushree Shah
2026-08-22 6:36 ` sashiko-bot
4 siblings, 1 reply; 11+ messages in thread
From: Tanushree Shah @ 2026-08-22 6:27 UTC (permalink / raw)
To: acme, jolsa, adrian.hunter, vmolnaro, mpetlan, tmricht, maddy,
irogers, namhyung
Cc: linux-perf-users, linuxppc-dev, atrajeev, hbathini, Tejas.Manhas1,
Tanushree.Shah, rostedt, Tanushree Shah
Add a shell test covering perf data convert --to-trace-dat,
alongside the existing --to-json and --to-ctf tests.
The test:
- skips if perf isn't linked with libtraceevent
- converts a plain sched:sched_switch recording to trace.dat
- converts a pipe-mode recording (perf record -o - | perf data
convert -i -)
- converts a recording with both tracepoint and non-tracepoint
events (sched:sched_switch,cycles)
- validates the resulting file with trace-cmd report if trace-cmd
is installed, otherwise falls back to a basic non-empty-file check
Signed-off-by: Tanushree Shah <tshah@linux.ibm.com>
---
...rf_data_converter_tracepoints_trace_dat.sh | 173 ++++++++++++++++++
1 file changed, 173 insertions(+)
create mode 100755 tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh
diff --git a/tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh b/tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh
new file mode 100755
index 000000000000..1c0bce270c38
--- /dev/null
+++ b/tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh
@@ -0,0 +1,173 @@
+#!/bin/bash
+# 'perf data convert --to-trace-dat' command test
+# SPDX-License-Identifier: GPL-2.0-or-later
+#
+# Copyright 2026, IBM Corporation
+# Author: Tanushree Shah <tshah@linux.ibm.com>
+
+set -e
+
+err=0
+
+perfdata=$(mktemp /tmp/__perf_test.perf.data.XXXXX)
+
+cleanup()
+{
+ rm -f "${perfdata}"
+ rm -f "${result}"
+ trap - exit term int
+}
+
+trap_cleanup()
+{
+ local exit_code=$?
+ if [ $exit_code -ne 0 ] && [ $exit_code -ne 2 ]; then
+ echo "Unexpected signal in ${FUNCNAME[1]}"
+ cleanup
+ exit 1
+ fi
+ cleanup
+ exit $exit_code
+}
+trap trap_cleanup exit term int
+
+result=$(mktemp /tmp/__perf_test.output.trace.dat.XXXXX)
+
+# Check if libtraceevent support is available
+if ! perf check feature libtraceevent
+then
+ echo "perf not linked with libtraceevent, skipping test"
+ exit 2
+fi
+
+# Check if trace-cmd is available for validation
+have_trace_cmd=0
+if command -v trace-cmd >/dev/null 2>&1; then
+ have_trace_cmd=1
+fi
+
+check_sched_switch()
+{
+ if ! perf list | grep -q 'sched:sched_switch'
+ then
+ echo "sched:sched_switch tracepoint not available, skipping test"
+ exit 2
+ fi
+}
+
+test_trace_converter_command()
+{
+ echo "Testing Perf Data Conversion Command to trace.dat"
+
+ if ! perf record -e sched:sched_switch -a -o "$perfdata" -- sleep 0.1
+ then
+ echo "Failed to record perf data"
+ err=1
+ return
+ fi
+
+ if ! perf data convert --to-trace-dat "$result" --force -i "$perfdata"
+ then
+ echo "Perf Data Converter Command to trace.dat [FAILED]"
+ err=1
+ return
+ fi
+
+ if [ -f "$result" ] && [ -s "$result" ] ; then
+ echo "Perf Data Converter Command to trace.dat [SUCCESS]"
+ else
+ echo "Perf Data Converter Command to trace.dat [FAILED]"
+ err=1
+ fi
+}
+
+test_trace_converter_pipe()
+{
+ echo "Testing Perf Data Conversion Command to trace.dat (Pipe mode)"
+
+ rm -f "$result"
+
+ if ! perf record -e sched:sched_switch -a -o - -- sleep 0.1 | \
+ perf data convert --to-trace-dat "$result" --force -i -
+ then
+ echo "Perf Data Converter Command to trace.dat (Pipe mode) [FAILED]"
+ err=1
+ return
+ fi
+
+ if [ -f "$result" ] && [ -s "$result" ]; then
+ echo "Perf Data Converter Command to trace.dat (Pipe mode) [SUCCESS]"
+ else
+ echo "Perf Data Converter Command to trace.dat (Pipe mode) [FAILED]"
+ err=1
+ fi
+}
+
+test_trace_converter_mixed_events()
+{
+ echo "Testing Perf Data Conversion with tracepoint and non-tracepoint events"
+
+ rm -f "$result"
+
+ # Record both tracepoint and non-tracepoint events
+ if ! perf record -e sched:sched_switch,cpu-clock -a -o "$perfdata" -- sleep 0.1
+ then
+ echo "Failed to record perf data (mixed events)"
+ err=1
+ return
+ fi
+
+ if ! perf data convert --to-trace-dat "$result" --force -i "$perfdata"
+ then
+ echo "Perf Data Converter with mixed events [FAILED]"
+ err=1
+ return
+ fi
+
+ if [ -f "$result" ] && [ -s "$result" ] ; then
+ echo "Perf Data Converter with mixed events [SUCCESS]"
+ else
+ echo "Perf Data Converter with mixed events [FAILED]"
+ err=1
+ fi
+}
+
+validate_trace_format()
+{
+ echo "Validating Perf Data Converted trace.dat file"
+
+ if [ ! -f "$result" ] ; then
+ echo "File not found [FAILED]"
+ err=1
+ return
+ fi
+
+ # If trace-cmd is available, use it to validate the file
+ if [ "$have_trace_cmd" -eq 1 ] ; then
+ if trace-cmd report -i "$result" >/dev/null 2>&1 ; then
+ echo "trace-cmd can read the file [SUCCESS]"
+ else
+ echo "trace-cmd cannot read the file [FAILED]"
+ err=1
+ fi
+ else
+ # Without trace-cmd, just check file exists and has content
+ if [ -s "$result" ] ; then
+ echo "Output file exists and has content [SUCCESS]"
+ else
+ echo "Output file is empty [FAILED]"
+ err=1
+ fi
+ fi
+}
+
+check_sched_switch
+test_trace_converter_command
+validate_trace_format
+test_trace_converter_pipe
+validate_trace_format
+test_trace_converter_mixed_events
+validate_trace_format
+
+cleanup
+exit ${err}
--
2.47.1
^ permalink raw reply related [flat|nested] 11+ messages in thread
* Re: [RFC PATCH v4 5/5] perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat'
2026-08-22 6:27 ` [RFC PATCH v4 5/5] perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat' Tanushree Shah
@ 2026-08-22 6:36 ` sashiko-bot
0 siblings, 0 replies; 11+ messages in thread
From: sashiko-bot @ 2026-08-22 6:36 UTC (permalink / raw)
To: Tanushree Shah; +Cc: linux-perf-users
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
- [Medium] Insecure temporary file reuse creates a time-of-check to time-of-use (TOCTOU) race condition and symlink attack vulnerability.
- [Medium] The commit message claims to test the 'cycles' event, but the script uses 'cpu-clock'.
- [Medium] The script leaks a temporary file (`${perfdata}.old`) in `/tmp` because `perf record` automatically backs up the existing non-empty output file.
--
--- Patch [5]: [RFC PATCH v4 5/5] perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat' ---
commit 9c32c2021bfda9ef574a5486cbcea176af233faa
Author: Tanushree Shah <tshah@linux.ibm.com>
perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat'
Add a shell test covering perf data convert --to-trace-dat,
alongside the existing --to-json and --to-ctf tests.
The test:
- skips if perf isn't linked with libtraceevent
- converts a plain sched:sched_switch recording to trace.dat
- converts a pipe-mode recording (perf record -o - | perf data
convert -i -)
> - converts a recording with both tracepoint and non-tracepoint
> events (sched:sched_switch,cycles)
[Severity: Medium]
Is this description accurate? The commit message mentions the "cycles"
event, but the code later tests the "cpu-clock" event instead.
- validates the resulting file with trace-cmd report if trace-cmd
is installed, otherwise falls back to a basic non-empty-file check
Signed-off-by: Tanushree Shah <tshah@linux.ibm.com>
> diff --git a/tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh b/tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh
> new file mode 100755
> index 0000000000000..1c0bce270c389
> --- /dev/null
> +++ b/tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh
> @@ -0,0 +1,173 @@
> +#!/bin/bash
> +# 'perf data convert --to-trace-dat' command test
> +# SPDX-License-Identifier: GPL-2.0-or-later
> +#
> +# Copyright 2026, IBM Corporation
> +# Author: Tanushree Shah <tshah@linux.ibm.com>
> +
> +set -e
> +
> +err=0
> +
> +perfdata=$(mktemp /tmp/__perf_test.perf.data.XXXXX)
> +
> +cleanup()
> +{
> + rm -f "${perfdata}"
[Severity: Medium]
Does this code leak the backup file created by perf?
When perf record is called a second time in test_trace_converter_mixed_events()
with -o "$perfdata", it detects that the file already exists and has content.
perf record automatically renames the existing file to "${perfdata}.old" before
creating a new one.
Since cleanup() only removes "${perfdata}", does the .old file get left
behind in /tmp?
> + rm -f "${result}"
> + trap - exit term int
> +}
> +
[ ... ]
> +test_trace_converter_pipe()
> +{
> + echo "Testing Perf Data Conversion Command to trace.dat (Pipe mode)"
> +
> + rm -f "$result"
> +
> + if ! perf record -e sched:sched_switch -a -o - -- sleep 0.1 | \
> + perf data convert --to-trace-dat "$result" --force -i -
[Severity: Medium]
Does this code create a time-of-check to time-of-use race condition?
The script securely creates a temporary file using mktemp, but explicitly
deletes it here with rm -f "$result".
Since the predictable and unsecured path is then immediately reused for
perf data convert, could a local attacker insert a symlink during this window,
causing perf data convert to overwrite arbitrary files?
> + then
> + echo "Perf Data Converter Command to trace.dat (Pipe mode) [FAILED]"
> + err=1
> + return
> + fi
> +
[ ... ]
> +test_trace_converter_mixed_events()
> +{
> + echo "Testing Perf Data Conversion with tracepoint and non-tracepoint events"
> +
> + rm -f "$result"
> +
> + # Record both tracepoint and non-tracepoint events
> + if ! perf record -e sched:sched_switch,cpu-clock -a -o "$perfdata" -- sleep 0.1
> + then
> + echo "Failed to record perf data (mixed events)"
> + err=1
> + return
> + fi
> +
> + if ! perf data convert --to-trace-dat "$result" --force -i "$perfdata"
[Severity: Medium]
Does this code repeat the symlink race condition mentioned above?
The securely created $result file is again explicitly deleted and its
path is reused here.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260822062729.882005-1-tshah@linux.ibm.com?part=5
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [RFC PATCH v4 3/5] perf data-convert: Add perf.data to trace.dat conversion backend
2026-08-22 6:27 ` [RFC PATCH v4 3/5] perf data-convert: Add perf.data to trace.dat conversion backend Tanushree Shah
@ 2026-08-22 6:38 ` sashiko-bot
0 siblings, 0 replies; 11+ messages in thread
From: sashiko-bot @ 2026-08-22 6:38 UTC (permalink / raw)
To: Tanushree Shah; +Cc: linux-perf-users
Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider:
- [High] Arbitrary file truncation and TOCTOU vulnerability via symlink in the `--force` path.
- [Medium] Directly accessing `perf_env` field without verifying initialization.
- [Low] Missing feature fallback stub for `trace_convert__perf2dat` in `data-convert.h`.
- [Low] Lost error code when `fopen()` fails in the `--force` path.
--
--- Patch [3]: [RFC PATCH v4 3/5] perf data-convert: Add perf.data to trace.dat conversion backend ---
commit aaf748bbf23e74fad4be33c3834c0413e8def0f8
Author: Tanushree Shah <tshah@linux.ibm.com>
perf data-convert: Add perf.data to trace.dat conversion backend
This commit adds a new conversion backend to perf to translate tracepoint
events from perf.data into trace.dat format. It leverages a new per-CPU
buffering mechanism and implements required callbacks for pipe mode.
> diff --git a/tools/perf/util/data-convert-trace.c b/tools/perf/util/data-convert-trace.c
> new file mode 100644
> index 0000000000000..445479fae8884
> --- /dev/null
> +++ b/tools/perf/util/data-convert-trace.c
[ ... ]
> + /*
> + * In pipe mode, CPU count and page size arrive via feature/tracing_data
> + * records before the first sample; initialize buffers lazily on first sample.
> + */
> + if (!cpu_buffers_initialized) {
> + int nr_cpus = trace_dat_session->header.env.nr_cpus_online;
[Severity: Medium]
Is it safe to access nr_cpus_online directly here without first verifying the
environment is initialized?
The perf subsystem guidelines suggest that accessing perf_env fields without
first verifying those fields are populated can lead to bugs.
[ ... ]
> + if (!opts->force) {
> + int fd = open(to_trace, O_WRONLY | O_CREAT | O_EXCL, 0644);
[ ... ]
> + } else {
> + trace_dat_fp = fopen(to_trace, "wb");
[Severity: High]
Could this fopen() in the --force path be susceptible to an arbitrary file
truncation symlink attack?
If an attacker creates a symlink at the destination path, fopen() with "wb"
will follow it and truncate the target file. If the conversion subsequently
fails, the cleanup routine calls unlink(to_trace), which removes the symlink
itself. This leaves the target file truncated while hiding the attacker's
symlink from discovery.
> + if (!trace_dat_fp) {
> + pr_err("Failed to open output file: %s\n", to_trace);
> + return -EINVAL;
[Severity: Low]
Should this return the actual -errno from fopen() instead of a hardcoded
-EINVAL so the exact filesystem error is not lost?
> + }
> + }
[ ... ]
> diff --git a/tools/perf/util/data-convert.h b/tools/perf/util/data-convert.h
> index a96240f156712..f041c23252263 100644
> --- a/tools/perf/util/data-convert.h
> +++ b/tools/perf/util/data-convert.h
[ ... ]
> +#ifdef HAVE_LIBTRACEEVENT
> +int trace_convert__perf2dat(const char *input, const char *to_trace,
> + struct perf_data_convert_opts *opts);
> +#endif /* HAVE_LIBTRACEEVENT */
[Severity: Low]
Does this declaration need an #else block with a compatible dummy inline stub?
Without a fallback stub when HAVE_LIBTRACEEVENT is absent, callers might
be forced to clutter their code with preprocessor blocks to prevent build
failures.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260822062729.882005-1-tshah@linux.ibm.com?part=3
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [RFC PATCH v4 1/5] perf/trace-dat: Add trace.dat export infrastructure
2026-08-22 6:27 ` [RFC PATCH v4 1/5] perf/trace-dat: Add trace.dat export infrastructure Tanushree Shah
@ 2026-08-22 6:40 ` sashiko-bot
0 siblings, 0 replies; 11+ messages in thread
From: sashiko-bot @ 2026-08-22 6:40 UTC (permalink / raw)
To: Tanushree Shah; +Cc: linux-perf-users
Thank you for your contribution! Sashiko AI review found 2 potential issue(s) to consider:
- [Medium] Large trace events encode their payload length incorrectly, causing alignment errors and event truncation.
- [Medium] Querying the local host system's `tracefs` for the `trace_clock` during file conversion leaks host environment state into the converted trace.
--
--- Patch [1]: [RFC PATCH v4 1/5] perf/trace-dat: Add trace.dat export infrastructure ---
commit 162f3240dcf1a2672aa00440674197558e8f2cec
Author: Tanushree Shah <tshah@linux.ibm.com>
perf/trace-dat: Add trace.dat export infrastructure
Add new utility files util/trace-dat.c and util/trace-dat.h implementing
the infrastructure for exporting perf.data tracepoints to trace.dat format
compatible with trace-cmd and KernelShark.
> diff --git a/tools/perf/util/trace-dat.c b/tools/perf/util/trace-dat.c
> index 0000000000000..55a7bd982c6f8
> --- /dev/null
> +++ b/tools/perf/util/trace-dat.c
[ ... ]
> +static int trace_dat__write_cpu_dat(FILE *fp, int cpu, unsigned long long *file_offset_out)
> +{
[ ... ]
> + memcpy(data_rec, &hdr_word, TRACE_DAT_WORD_SIZE);
> +
> + /* Large events: write actual byte length after header */
> + if (type_len == 0)
> + memcpy(data_rec + TRACE_DAT_WORD_SIZE, &data_len, TRACE_DAT_WORD_SIZE);
[Severity: Medium]
When writing large events (where type_len == 0), is `data_len` the correct
value to write here? Ftrace ring buffer parsers typically expect this field
to contain the size of the payload plus the 4-byte size field itself,
properly padded to 4-byte alignment, rather than just the raw event size.
Could this lead to event truncation or alignment errors when parsing?
[ ... ]
> +int trace_dat__write_options_section1(void)
> +{
[ ... ]
> + /* TRACECLOCK option */
> + opt_id = TRACE_DAT_OPTION_TRACECLOCK;
> +
> + path = get_tracing_file("trace_clock");
> + if (path) {
> + clock_file = fopen(path, "r");
> + put_tracing_file(path);
> + } else {
> + clock_file = NULL;
> + }
[Severity: Medium]
Does querying the host system's tracefs mount here break cross-platform
analysis? Since this is an offline conversion of a perf.data file, the host
machine's current trace clock might not match the clock used when the trace
was originally recorded. Should this extract the clock from the perf_env
instead?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260822062729.882005-1-tshah@linux.ibm.com?part=1
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [RFC PATCH v4 4/5] perf data: Add --to-trace-dat option for converting perf.data tracepoint events into trace.dat format
2026-08-22 6:27 ` [RFC PATCH v4 4/5] perf data: Add --to-trace-dat option for converting perf.data tracepoint events into trace.dat format Tanushree Shah
@ 2026-08-22 6:44 ` sashiko-bot
0 siblings, 0 replies; 11+ messages in thread
From: sashiko-bot @ 2026-08-22 6:44 UTC (permalink / raw)
To: Tanushree Shah; +Cc: linux-perf-users
Thank you for your contribution! Sashiko AI review found 6 potential issue(s) to consider:
- [High] Trace parser desynchronization: Large event array lengths must be encoded as the fully padded size minus the length word.
- [High] Out of bounds rejection: CPU array limits must be scoped to `nr_cpus_avail` rather than `nr_cpus_online`.
- [High] Silent I/O truncation: Stream checks must explicitly assert identical element counts instead of binary negation.
- [High] Architecture violation: Page header offsets must dynamically query the ftrace event format to accommodate 32-bit `local_t` structures.
- [Low] Build architecture violation: The `data-convert.h` header must include a dummy inline fallback stub for `trace_convert__perf2dat`.
- [Low] Dead code elimination: The `skipped` sample counter is incremented but structurally isolated and unread.
--
--- Patch [4]: [RFC PATCH v4 4/5] perf data: Add --to-trace-dat option for converting perf.data tracepoint events into trace.dat format ---
commit 9a258038767861211db47cd6d942b553034516f1
Author: Tanushree Shah <tshah@linux.ibm.com>
perf data: Add --to-trace-dat option for converting perf.data tracepoint events into trace.dat format
Add new command-line option to perf data convert for generating
trace.dat output files.
> diff --git a/tools/perf/builtin-data.c b/tools/perf/builtin-data.c
> index 1dd73ed6bdcb8..c9c863197d02c 100644
> --- a/tools/perf/builtin-data.c
> +++ b/tools/perf/builtin-data.c
[ ... ]
> @@ -63,10 +72,44 @@ static int cmd_data_convert(int argc, const char **argv)
[ ... ]
> +#ifdef HAVE_LIBTRACEEVENT
> + if (trace_dat_output)
> + return trace_convert__perf2dat(input_name ? input_name : "perf.data",
> + trace_dat_output, &opts);
> +#endif
[Severity: Low]
Does the data-convert.h header need a dummy inline fallback stub for
trace_convert__perf2dat()?
Without an #else block providing an inline dummy stub, could this cause a
build failure when compiling without HAVE_LIBTRACEEVENT?
> diff --git a/tools/perf/util/data-convert-trace.c b/tools/perf/util/data-convert-trace.c
> index 445479fae8884..425b3aaa3f026 100644
> --- a/tools/perf/util/data-convert-trace.c
> +++ b/tools/perf/util/data-convert-trace.c
> @@ -31,6 +32,10 @@
> struct trace_convert {
> struct perf_tool tool;
> u64 events_count;
> + struct perf_time_interval *ptime_range;
> + int range_size;
> + int range_num;
> + u64 skipped;
> };
[Severity: Low]
Is this skipped counter ever read?
It is incremented when events are outside the specified time range, but
appears to be structurally isolated and never emitted in the final
conversion summary.
> @@ -86,6 +91,11 @@ static int process_sample_event(const struct perf_tool *tool,
> cpu_buffers_initialized = true;
> }
>
> + if (perf_time__ranges_skip_sample(tc->ptime_range, tc->range_num, sample->time)) {
> + tc->skipped++;
> + return 0;
> + }
> +
[Severity: High]
Are valid samples from sparse CPU topologies rejected here?
In process_sample_event(), the CPU buffers are initialized using
nr_cpus_online:
int nr_cpus = trace_dat_session->header.env.nr_cpus_online;
But trace_dat__collect_cpu_event() rejects samples if the CPU index is >=
trace_dat_nr_cpus:
if (!trace_cpu_data || cpu < 0 || cpu >= trace_dat_nr_cpus)
return -EINVAL;
If a system has hotplugged CPUs (e.g. CPUs 0 and 127 are online, making
nr_cpus_online 2), a valid sample from CPU 127 will be evaluated against
cpu >= 2. This causes trace_dat__collect_cpu_event() to return -EINVAL,
aborting the conversion.
Should this use nr_cpus_avail instead of nr_cpus_online?
> diff --git a/tools/perf/util/trace-dat.c b/tools/perf/util/trace-dat.c
> index f71e03716e27e..1b236745bcacc 100644
> --- a/tools/perf/util/trace-dat.c
> +++ b/tools/perf/util/trace-dat.c
[ ... ]
> @@ -288,7 +292,7 @@ static int trace_dat__write_cpu_dat(FILE *fp, struct tep_handle *pevent,
> /* Check page fit BEFORE allocating data record */
> if (page_size_used + needed_size >
> trace_dat_page_size - TRACE_DAT_RECORD_HEADER_SIZE) {
> - ret = trace_dat__write_page(fp, pevent, base_ts,
> + ret = trace_dat__write_page(fp, pevent, page_base_ts,
> page_records, page_rec_sizes,
> nr_page_recs);
[Severity: High]
Could the length word for large trace events be miscalculated?
Later in trace_dat__write_cpu_dat(), the array length is encoded:
/* Large events: write actual byte length after header */
if (type_len == 0) {
unsigned int data_len_out = to_file_u32(pevent, data_len);
memcpy(data_rec + TRACE_DAT_WORD_SIZE, &data_len_out, TRACE_DAT_WORD_SIZE);
}
For large events (>=29 words), the ftrace ring buffer format dictates that
the first word of the payload contains the length of the data plus the
length word itself (data_len + 4).
By writing just data_len, might this cause downstream tools like trace-cmd
to compute the event size incorrectly and lose parser synchronization?
[Severity: High]
Can a short write silently corrupt the trace.dat file?
In trace_dat__write_page(), the file write checks for a boolean false:
if (!fwrite(page, 1, trace_dat_page_size, fp)) {
free(page);
return -EIO;
}
Since fwrite() returns the number of elements written, a short write (e.g.
512 out of 4096 bytes) evaluates to true.
Will this silently drop data without reporting an error to the user if the
disk runs out of space? Should the stream check explicitly assert
fwrite(...) == trace_dat_page_size?
[Severity: High]
Will this hardcoded page header size break parsing for trace data recorded
on 32-bit architectures?
In trace_dat__write_page(), the page header offsets are hardcoded:
/* Byte-swap page header for cross-arch compatibility */
ts_out = to_file_u64(pevent, base_ts);
commit_out = to_file_u64(pevent, commit);
memcpy(page, &ts_out, sizeof(ts_out));
memcpy(page + sizeof(ts_out), &commit_out, sizeof(commit_out));
The ftrace page header includes a local_t commit field, which is 8 bytes on
64-bit systems but 4 bytes on 32-bit systems (making the header 12 bytes).
Because this unconditionally assumes a 16-byte header and writes 8 bytes for
commit_out, trace-cmd reading 32-bit metadata will likely parse the remaining
4 bytes as the first trace event.
Should the page header offsets and sizing be dynamically queried from the
ftrace event format?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260822062729.882005-1-tshah@linux.ibm.com?part=4
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [RFC PATCH v4 2/5] perf/trace-event: Write trace.dat metadata sections during parsing
2026-08-22 6:27 ` [RFC PATCH v4 2/5] perf/trace-event: Write trace.dat metadata sections during parsing Tanushree Shah
@ 2026-08-22 6:44 ` sashiko-bot
0 siblings, 0 replies; 11+ messages in thread
From: sashiko-bot @ 2026-08-22 6:44 UTC (permalink / raw)
To: Tanushree Shah; +Cc: linux-perf-users
Thank you for your contribution! Sashiko AI review found 5 potential issue(s) to consider:
New issues:
- [Critical] Missing bounds check on `raw_size` leads to massive out-of-bounds read and heap buffer overflow.
- [High] Hardcoded 16-byte page header breaks parsing of trace data recorded on 32-bit architectures.
- [Medium] Host state leakage / Cross-platform analysis violation when embedding the TRACECLOCK option.
- [Medium] Signed integer overflow causing undefined behavior in TIME_EXTEND bitshift.
Pre-existing issues:
- [Critical] Integer overflow in `read_saved_cmdline` size calculation leads to a heap buffer overflow on 32-bit systems.
--
--- Patch [2]: [RFC PATCH v4 2/5] perf/trace-event: Write trace.dat metadata sections during parsing ---
commit 75d79e7a5c5c4a1484c389705c19fc30dbdd1deb
Author: Tanushree Shah <tshah@linux.ibm.com>
perf/trace-event: Write trace.dat metadata sections during parsing
Perf already captures the tracing metadata as a part of
data section in perf.data
When trace_dat_fp is set, write trace.dat compatible metadata
sections using the perf provided raw buffers.
> diff --git a/tools/perf/util/trace-dat.c b/tools/perf/util/trace-dat.c
> --- a/tools/perf/util/trace-dat.c
> +++ b/tools/perf/util/trace-dat.c
[ ... ]
> @@ -161,8 +164,12 @@ static int trace_dat__write_page(FILE *fp, unsigned long long base_ts,
> commit += rec_sizes[i];
> }
>
> - memcpy(page, &base_ts, sizeof(base_ts));
> - memcpy(page + sizeof(base_ts), &commit, sizeof(commit));
> + /* Byte-swap page header for cross-arch compatibility */
> + ts_out = to_file_u64(pevent, base_ts);
> + commit_out = to_file_u64(pevent, commit);
> +
> + memcpy(page, &ts_out, sizeof(ts_out));
> + memcpy(page + sizeof(ts_out), &commit_out, sizeof(commit_out));
[Severity: High]
Does this hardcode a 16-byte page header regardless of the original
architecture's commit size?
If the original trace was recorded on a 32-bit machine, its header_page
metadata usually specifies the commit field as a 4-byte local_t. Writing a
hardcoded 8-byte commit field here might cause parsers to misalign by 4
bytes, potentially interpreting the upper half of the 8-byte commit as
the first event record's header and breaking parsing.
[ ... ]
> @@ -229,11 +237,18 @@ static int trace_dat__write_cpu_dat(FILE *fp, int cpu, unsigned long long *file_
> if (!extend)
> return -ENOMEM;
>
> - extend_hdr =
> - ((time_delta & TRACE_DAT_RECORD_TIME_MASK) <<
> - TRACE_DAT_RECORD_TIME_SHIFT) |
> - TRACE_DAT_RECORD_TYPE_TIME_EXTEND;
> - delta_upper = time_delta >> TRACE_DAT_RECORD_TIME_SHIFT;
> + if (tep_is_file_bigendian(pevent)) {
> + extend_hdr = (time_delta & TRACE_DAT_RECORD_TIME_MASK) |
> + (TRACE_DAT_RECORD_TYPE_TIME_EXTEND << 27);
> + delta_upper = time_delta >> 27;
[Severity: Medium]
Could this left shift result in undefined behavior?
Assuming TRACE_DAT_RECORD_TYPE_TIME_EXTEND is defined as 30 (a 32-bit signed
integer literal), shifting it left by 27 bits yields 4026531840. This exceeds
the maximum representable value of a 32-bit signed integer, which can cause
undefined behavior in C.
[ ... ]
> @@ -311,8 +332,11 @@ static int trace_dat__write_cpu_dat(FILE *fp, int cpu, unsigned long long *file_
> memcpy(data_rec, &hdr_word, TRACE_DAT_WORD_SIZE);
>
> /* Large events: write actual byte length after header */
> - if (type_len == 0)
> - memcpy(data_rec + TRACE_DAT_WORD_SIZE, &data_len, TRACE_DAT_WORD_SIZE);
> + if (type_len == 0) {
> + unsigned int data_len_out = to_file_u32(pevent, data_len);
> +
> + memcpy(data_rec + TRACE_DAT_WORD_SIZE, &data_len_out, TRACE_DAT_WORD_SIZE);
> + }
>
> memcpy(data_rec + payload_offset, event->raw, data_len);
[Severity: Critical]
Is it possible for a crafted perf.data file to cause an out-of-bounds read
and heap buffer overflow here?
If the original sample's raw_size is maliciously set near UINT_MAX
(e.g., 0xFFFFFFFF), the data_rec_size calculation in
trace_dat__write_cpu_dat() calculates data_rec_size = 8 + data_len, which
wraps around to 7. This bypasses the page size check and allocates an 8-byte
heap buffer. The memcpy() here then attempts to copy gigabytes of data into
that small buffer.
Similarly, earlier in trace_dat__collect_cpu_event(), malloc(0xFFFFFFFF)
could succeed on a 64-bit host, causing the subsequent memcpy() to read
gigabytes out-of-bounds.
[ ... ]
> @@ -464,16 +500,17 @@ int trace_dat__write_options_section1(void)
> payload_start = ftell(trace_dat_fp);
>
> /* CPUCOUNT option */
[ ... ]
> /* TRACECLOCK option */
> - opt_id = TRACE_DAT_OPTION_TRACECLOCK;
> + opt_id = to_file_u16(pevent, TRACE_DAT_OPTION_TRACECLOCK);
>
> path = get_tracing_file("trace_clock");
> if (path) {
> clock_file = fopen(path, "r");
[Severity: Medium]
Does this inadvertently leak the conversion host's state into the
converted trace data?
Since this dynamically queries the local host's trace_clock via
get_tracing_file(), it embeds the local machine's unparsed clock file contents
(such as "[local] global...") instead of preserving the original recorded
machine's clock metadata, potentially breaking cross-platform reproducibility.
[ ... ]
> diff --git a/tools/perf/util/trace-event-read.c b/tools/perf/util/trace-event-read.c
> --- a/tools/perf/util/trace-event-read.c
> +++ b/tools/perf/util/trace-event-read.c
> @@ -383,6 +615,27 @@ static int read_saved_cmdline(struct tep_handle *pevent)
> pr_debug("error reading saved cmdlines\n");
> goto out;
> }
> + /* Write cmdlines section with data */
> + if (trace_dat_fp && !trace_dat_write_failed) {
> + unsigned short section_id = to_file_u16(pevent, TRACE_DAT_SECTION_CMDLINE);
> + unsigned short flags = to_file_u16(pevent, 0);
[Severity: Critical]
This is a pre-existing issue, but does the size calculation for the
saved command line buffer overflow on 32-bit systems?
Looking at read_saved_cmdline(), when size is read as a 64-bit value, if it
is 0xFFFFFFFF, the allocation calculation of size + 1 overflows to
0x100000000.
On 32-bit architectures, this is truncated to 0, causing malloc(0) to return
a small valid pointer. The subsequent do_read(buf, size) then passes
0xFFFFFFFF to the read() syscall, which can overwrite the heap and
potentially cause a denial of service or arbitrary code execution.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260822062729.882005-1-tshah@linux.ibm.com?part=2
^ permalink raw reply [flat|nested] 11+ messages in thread
end of thread, other threads:[~2026-08-22 6:44 UTC | newest]
Thread overview: 11+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-08-22 6:27 [RFC PATCH v4 0/5] Add perf.data tracepoint events to trace.dat conversion Tanushree Shah
2026-08-22 6:27 ` [RFC PATCH v4 1/5] perf/trace-dat: Add trace.dat export infrastructure Tanushree Shah
2026-08-22 6:40 ` sashiko-bot
2026-08-22 6:27 ` [RFC PATCH v4 2/5] perf/trace-event: Write trace.dat metadata sections during parsing Tanushree Shah
2026-08-22 6:44 ` sashiko-bot
2026-08-22 6:27 ` [RFC PATCH v4 3/5] perf data-convert: Add perf.data to trace.dat conversion backend Tanushree Shah
2026-08-22 6:38 ` sashiko-bot
2026-08-22 6:27 ` [RFC PATCH v4 4/5] perf data: Add --to-trace-dat option for converting perf.data tracepoint events into trace.dat format Tanushree Shah
2026-08-22 6:44 ` sashiko-bot
2026-08-22 6:27 ` [RFC PATCH v4 5/5] perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat' Tanushree Shah
2026-08-22 6:36 ` sashiko-bot
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox