* [RFC PATCH v3 0/5] Add perf.data tracepoint events to trace.dat conversion
@ 2026-08-03 14:59 Tanushree Shah
2026-08-03 14:59 ` [RFC PATCH v3 1/5] perf/trace-dat: Add trace.dat export infrastructure Tanushree Shah
` (5 more replies)
0 siblings, 6 replies; 14+ messages in thread
From: Tanushree Shah @ 2026-08-03 14:59 UTC (permalink / raw)
To: acme, jolsa, adrian.hunter, vmolnaro, mpetlan, tmricht, maddy,
irogers, namhyung
Cc: linux-perf-users, linuxppc-dev, atrajeev, hbathini, Tejas.Manhas1,
Tanushree.Shah, rostedt, Tanushree Shah
This RFC patch series introduces support for converting perf.data files
containing tracepoint events into trace.dat format, enabling seamless
visualization and analysis using KernelShark.
======================
Background and Motivation
======================
Currently, perf and trace-cmd operate as separate tracing ecosystems with
incompatible data formats. Users who collect tracepoint data with
'perf record' cannot easily visualize it in KernelShark's graphical
timeline view or leverage trace-cmd's analysis capabilities.
This creates workflow friction when users need to:
- Visualize perf tracepoint data in KernelShark's interactive graphical
timeline
- Share trace data between perf and trace-cmd workflows and toolchains
- Perform architecture-independent conversion and analysis of traces
This conversion bridge eliminates these barriers by enabling seamless
data exchange between perf and trace-cmd ecosystems, allowing users to
choose the best tool for each analysis phase.
======================
Implementation Overview
======================
The series implements the trace.dat file format specification (version 7)
within perf's data conversion framework.
**Patch 1/5: Core trace.dat Export Infrastructure**
Introduces util/trace-dat.c and util/trace-dat.h implementing:
- Per-CPU raw event buffer management (init, collect, free)
- Ftrace ring buffer page construction
- trace.dat section writers (strings, options, flyrecord sections)
**Patch 2/5: Metadata Integration**
Extends util/trace-event-read.c to write trace.dat metadata during
perf.data
parsing:
- Initial format header (magic, version, endian, page size, compression)
- Section 16: HEADER INFO (header_page + header_event)
- Section 17: FTRACE EVENT FORMATS
- Section 18: EVENT FORMATS (per system/event format files)
- Section 19: KALLSYMS
- Section 21: CMDLINES
- Section 15: STRINGS (written last after all sections)
**Patch 3/5: Conversion Backend**
Implements util/data-convert-trace.c with trace_convert__perf2dat()
function:
- Processes PERF_TYPE_TRACEPOINT samples via process_sample_event()
- Collects raw event data per-CPU using trace_dat__collect_cpu_event()
- Writes OPTIONS sections (CPUCOUNT, TRACECLOCK, metadata offsets)
- Writes FLYRECORD section with per-CPU ring buffer pages
**Patch 4/5: User Interface**
Extends tools/perf/builtin-data.c with --to-trace-dat option:
- Adds command-line option for trace.dat output
- Mutually exclusive with --to-ctf and --to-json
- Calls trace_convert__perf2dat() to perform conversion
**Patch 5/5: Shell Test**
Adds a shell test (tools/perf/tests/shell/) covering normal tracepoint
recordings, pipe mode and mixed tracepoint/non-tracepoint recording
conversions and --force flag behaviour.
======================
Current Implementation Details
======================
**trace.dat Format Version:**
The implementation currently targets trace.dat format version 7, which
is the stable version supported by current trace-cmd releases (v3.x).
This version is hardcoded to ensure compatibility with existing
trace-cmd and KernelShark installations. Future enhancements could add
version negotiation or support for newer format versions as they become
standardized.
**Compression Strategy:**
Compression is explicitly disabled (set to NONE) in the generated
trace.dat files.
This design choice:
- Simplifies the initial implementation and testing
- Ensures maximum compatibility across trace-cmd versions
- Avoids external compression library dependencies
Future work could add support for various compression algorithms (zlib,
zstd, lz4) with runtime selection via command-line options, significantly
reducing file sizes for large traces.
======================
Usage Example
======================
```bash
*Record tracepoint events with perf*
perf record -e sched:sched_switch -e sched:sched_wakeup -a sleep 10
*Convert to trace.dat format*
perf data convert --to-trace-dat=output.dat
*Verify trace.dat structure*
trace-cmd dump --summary output.dat
*Analyze with trace-cmd*
trace-cmd report output.dat
*Visualize in KernelShark*
kernelshark output.dat
```
**Conversion Output:**
```
[ perf data convert: Converted 'perf.data' into trace.dat format
'output.dat' ]
[ perf data convert: Converted 2684 events ]
```
**trace-cmd dump --summary Output:**
```
Tracing meta data in file output.dat:
[Initial format]
7 [Version]
0 [Little endian]
8 [Bytes in a long]
65536 [Page size, bytes]
none [Compression algorithm]
[Compression version]
[buffer "", "local" clock, 65536 page size, 16 cpus, 1048576 bytes
flyrecord data]
[10 options]
[Saved command lines, 0 bytes]
[Kallsyms, 0 bytes]
[Ftrace format, 0 events]
[Header page, 206 bytes]
[Header event, 205 bytes]
[Events format, 1 systems]
[9 sections]
```
======================
Testing and Verification
======================
The series has been extensively tested with:
- Various tracepoint events (sched, irq, syscalls, block I/O)
- Mixed recordings containing both tracepoint and non-tracepoint events
only tracepoints converted)
- Verification with trace-cmd report and KernelShark visualization
- Memory leak testing with Valgrind (0 bytes leaked).
- Cross-architecture testing: v1 tested x86_64 and ppc64le. v2 adds
s390 (big-endian) perf.data converted on both ppc64le
(little-endian) and x86_64 (little-endian) hosts, in addition to
same-arch x86_64 (LE->LE) and ppc64le (BE->BE) conversion.
- Pipe mode support has been tested end-to-end. (in v2)
All generated trace.dat files successfully open in:
- trace-cmd report (v3.1+)
- KernelShark (v2.0+)
======================
Next Steps
======================
We would highly appreciate reviews, comments, and feedback on:
- The overall architectural approach and integration points
- Compatibility considerations with trace-cmd ecosystem
- Performance characteristics for large-scale traces
- Additional use cases or workflow scenarios
- Future enhancement priorities
---
Changes in v3
- Rebase on latest perf-tools-next and resolve merge conflicts.
- Drop evsel parameter from process_sample_event() following upstream
commit "perf tool: Remove evsel from tool APIs that pass the sample".
Changes in v2
Addressing the Sashiko AI review findings on v1:
Cross-arch correctness:
- Introduce to_file_u16/u32/u64 helpers (wrapping tep_read_number())
to write all multi-byte fields in the recorded machine's byte order;
apply throughout metadata sections and flyrecord page/record headers
(ts, commit, TIME_EXTEND, large-event data_len).
- Fix flyrecord record header bit layout for big-endian files.
The record header word bit layout differs by file endianness,
matching kbuffer-parse.c type_len4host()/ts4host():
LE: type_len in bits [4:0], time_delta in bits [31:5]
BE: type_len in bits [31:27], time_delta in bits [26:0]
Pipe mode:
- Add process_attr(), process_feature(), and process_tracing_data()
callbacks required for pipe mode operation.
- Defer CPU buffer initialisation until the first tracepoint sample,
after process_feature()/process_tracing_data() have populated the
session header. This ensures the recorded machine's CPU count is
used rather than the host's - critical for cross-platform analysis.
Format compliance:
- Implement TIME_EXTEND records for timestamp deltas >27 bits to
prevent silent truncation and maintain chronological ordering.
- Fix large event encoding (>=29 words): use type_len=0 with a
separate 32-bit length word, avoiding collision with reserved types
(PADDING=29, TIME_EXTEND=30, TIME_STAMP=31).
- Add bounds check rejecting records larger than a page payload before
batching, preventing heap overflow in trace_dat__write_page().
- Fix flyrecord section_size to exclude the 16-byte section header,
matching trace.dat specification and trace-cmd behaviour.
CLI behavior:
- Fix --force flag: open with O_CREAT|O_EXCL when force is not set,
failing with -EEXIST instead of silently overwriting existing files.
Memory safety:
- Fix realloc overwrite of cpu_events->events and page_records on
failure: use temporary pointers, only commit on success.
- Fix use-after-free/double-free in sequential page_records realloc
failure: replace with malloc+memcpy+free pattern.
- Fix section_size computed from before section header position.
- Add NULL checks for get_tracing_file(), calloc() padding, and
trace_dat_options_offset assignment on write failure.
- Use goto out_free on record allocation failure to avoid leaking
accumulated page_records entries.
- Replace direct read() with do_read() in read_proc_kallsyms() to
handle short reads correctly.
- On fwrite failure, set trace_dat_write_failed and continue parsing
so that perf.data processing completes normally.
Testing (new in v2):
- Add shell test covering conversion, trace-cmd dump validation,
sched_switch event verification, and --force flag behaviour.
Documentation (new in v2):
- Add documentation for 'perf data convert --to-trace-dat', covering
usage and supported options.
v1: https://lore.kernel.org/linux-perf-users/20260608125951.90425-2-tshah@linux.ibm.com/
Tanushree Shah (5):
perf/trace-dat: Add trace.dat export infrastructure
perf/trace-event: Write trace.dat metadata sections during parsing
perf data-convert: Add perf.data to trace.dat conversion backend
perf data: Add --to-trace-dat option for converting perf.data
tracepoint events into trace.dat format
perf test: Add test validating trace.dat generated by 'perf data
convert --to-trace-dat'
tools/perf/Documentation/perf-data.txt | 7 +
tools/perf/builtin-data.c | 43 +
...rf_data_converter_tracepoints_trace_dat.sh | 169 ++++
tools/perf/util/Build | 2 +
tools/perf/util/data-convert-trace.c | 241 +++++
tools/perf/util/data-convert.h | 4 +
tools/perf/util/trace-dat.c | 876 ++++++++++++++++++
tools/perf/util/trace-dat.h | 113 +++
tools/perf/util/trace-event-read.c | 307 +++++-
9 files changed, 1755 insertions(+), 7 deletions(-)
create mode 100755 tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh
create mode 100644 tools/perf/util/data-convert-trace.c
create mode 100644 tools/perf/util/trace-dat.c
create mode 100644 tools/perf/util/trace-dat.h
--
2.47.1
^ permalink raw reply [flat|nested] 14+ messages in thread* [RFC PATCH v3 1/5] perf/trace-dat: Add trace.dat export infrastructure 2026-08-03 14:59 [RFC PATCH v3 0/5] Add perf.data tracepoint events to trace.dat conversion Tanushree Shah @ 2026-08-03 14:59 ` Tanushree Shah 2026-08-03 15:14 ` sashiko-bot 2026-08-03 14:59 ` [RFC PATCH v3 2/5] perf/trace-event: Write trace.dat metadata sections during parsing Tanushree Shah ` (4 subsequent siblings) 5 siblings, 1 reply; 14+ messages in thread From: Tanushree Shah @ 2026-08-03 14:59 UTC (permalink / raw) To: acme, jolsa, adrian.hunter, vmolnaro, mpetlan, tmricht, maddy, irogers, namhyung Cc: linux-perf-users, linuxppc-dev, atrajeev, hbathini, Tejas.Manhas1, Tanushree.Shah, rostedt, Tanushree Shah Add new utility files util/trace-dat.c and util/trace-dat.h implementing the infrastructure for exporting perf.data tracepoints to trace.dat format compatible with trace-cmd and KernelShark. trace-dat.c defines all globals and functions needed for: - Per-cpu raw event buffer management (init_cpu_buffers, collect_cpu_event, free_cpu_buffers). - ftrace ring buffer page construction (write_page, write_cpu_dat), including TIME_EXTEND records for large timestamp deltas. - trace.dat section writers (write_strings_section, write_options_section1, write_options_section2, write_flyrecord_section). trace-dat.h declares all globals and function prototypes to be used by data-convert-trace.c and trace-event-read.c. Signed-off-by: Tanushree Shah <tshah@linux.ibm.com> --- tools/perf/util/Build | 1 + tools/perf/util/trace-dat.c | 825 ++++++++++++++++++++++++++++++++++++ tools/perf/util/trace-dat.h | 83 ++++ 3 files changed, 909 insertions(+) create mode 100644 tools/perf/util/trace-dat.c create mode 100644 tools/perf/util/trace-dat.h diff --git a/tools/perf/util/Build b/tools/perf/util/Build index 330311cac550..246d05ce25cd 100644 --- a/tools/perf/util/Build +++ b/tools/perf/util/Build @@ -98,6 +98,7 @@ perf-util-y += trace-event-scripting.o perf-util-$(CONFIG_LIBTRACEEVENT) += trace-event.o perf-util-$(CONFIG_LIBTRACEEVENT) += trace-event-parse.o perf-util-$(CONFIG_LIBTRACEEVENT) += trace-event-read.o +perf-util-$(CONFIG_LIBTRACEEVENT) += trace-dat.o perf-util-y += sort.o perf-util-y += hist.o perf-util-y += util.o diff --git a/tools/perf/util/trace-dat.c b/tools/perf/util/trace-dat.c new file mode 100644 index 000000000000..55a7bd982c6f --- /dev/null +++ b/tools/perf/util/trace-dat.c @@ -0,0 +1,825 @@ +// SPDX-License-Identifier: GPL-2.0-or-later +/* + * Copyright 2026, IBM Corporation + * Author: Tanushree Shah <tshah@linux.ibm.com> + * + * trace-dat.c + * + * This file implements the trace.dat format writer for perf tool. + * It collects trace events from multiple CPUs and writes them in + * the trace-cmd compatible format. + */ +#include <stdio.h> +#include <stdlib.h> +#include <string.h> +#include <unistd.h> +#include <errno.h> +#include "api/fs/tracing_path.h" +#include "trace-dat.h" +#include "trace-event.h" +#include "session.h" +#include "header.h" +#include "../perf.h" +#include "debug.h" + +/* ftrace ring buffer constants for trace.dat flyrecord section + * + * Each page has a 16-byte header (timestamp + commit size), followed by + * variable-length records. Each record has a 4-byte header word encoding: + * Bits 0-4: Type/Length field (5 bits, masked by TYPE_LEN_MASK) + * Bits 5-31: Time delta from page base timestamp (27 bits, masked by TIME_MASK) + */ +#define TRACE_DAT_RECORD_HEADER_SIZE 16 /* Page header: 8-byte ts + 8-byte commit */ +#define TRACE_DAT_RECORD_TYPE_LEN_MASK 0x1F /* Extract lower 5 bits for type/length */ +#define TRACE_DAT_RECORD_TIME_SHIFT 5 /* Shift to extract time delta */ +#define TRACE_DAT_RECORD_TIME_MASK 0x07FFFFFF /* Mask for 27-bit time delta */ +#define TRACE_DAT_WORD_SIZE 4 /* Records aligned to 4-byte boundaries */ +#define TRACE_DAT_WORD_ALIGN_MASK 3 + +/* Initial capacity for per-CPU event buffer (grows by doubling) */ +#define INITIAL_EVENT_CAPACITY 1024 +/* Initial capacity for page record array (grows by doubling) */ +#define INITIAL_PAGE_RECORD_CAPACITY 64 +/* Buffer size for reading trace_clock string from debugfs/tracefs */ +#define CLOCK_BUFFER_SIZE 256 + +FILE *trace_dat_fp; +int trace_dat_page_size; +int trace_dat_nr_cpus; +long trace_dat_options_offset; +long trace_dat_header_info_offset; +long trace_dat_events_format_offset; +long trace_dat_ftrace_format_offset; +long trace_dat_kallsyms_offset; +long trace_dat_cmdline_offset; +long trace_dat_next_options_offset; + + +/** + * struct cpu_event - Single trace event from a CPU + * @ts: Timestamp of the event + * @raw: Raw event data + * @raw_size: Size of raw event data in bytes + */ +struct cpu_event { + unsigned long long ts; + void *raw; + unsigned int raw_size; +}; + +/** + * struct cpu_events - Collection of trace events for a single CPU + * @events: Array of events + * @count: Number of events currently stored + * @capacity: Maximum number of events that can be stored + */ +struct cpu_events { + struct cpu_event *events; + int count; + int capacity; +}; + +static struct cpu_events *trace_cpu_data; +static long *buffer_opt_cpu_offsets_pos; +static long opt_payload_start; + +/* Allocate per-cpu event buffers for tracepoint data collection */ +int trace_dat__init_cpu_buffers(int nr_cpus) +{ + trace_cpu_data = calloc(nr_cpus, sizeof(struct cpu_events)); + if (!trace_cpu_data) + return -ENOMEM; + buffer_opt_cpu_offsets_pos = calloc(nr_cpus, sizeof(long)); + if (!buffer_opt_cpu_offsets_pos) { + free(trace_cpu_data); + trace_cpu_data = NULL; + return -ENOMEM; + } + trace_dat_nr_cpus = nr_cpus; + return 0; +} + +/* Store raw tracepoint event data in per-cpu buffer for trace.dat + * flyrecord + */ +int trace_dat__collect_cpu_event(int cpu, unsigned long long ts, + void *raw, unsigned int raw_size) +{ + struct cpu_events *cpu_events; + void *raw_copy; + + if (!trace_cpu_data || cpu < 0 || cpu >= trace_dat_nr_cpus) + return -EINVAL; + + if (!raw || raw_size == 0) + return -EINVAL; + + cpu_events = &trace_cpu_data[cpu]; + + if (cpu_events->count >= cpu_events->capacity) { + int new_capacity = cpu_events->capacity ? + cpu_events->capacity * 2 : INITIAL_EVENT_CAPACITY; + struct cpu_event *tmp; + + tmp = realloc(cpu_events->events, new_capacity * sizeof(*cpu_events->events)); + if (!tmp) + return -ENOMEM; + cpu_events->events = tmp; + cpu_events->capacity = new_capacity; + } + + raw_copy = malloc(raw_size); + if (!raw_copy) + return -ENOMEM; + + memcpy(raw_copy, raw, raw_size); + cpu_events->events[cpu_events->count].ts = ts; + cpu_events->events[cpu_events->count].raw = raw_copy; + + cpu_events->events[cpu_events->count].raw_size = raw_size; + cpu_events->count++; + + return 0; +} + +/* Write a single page of trace records */ +static int trace_dat__write_page(FILE *fp, unsigned long long base_ts, + char **records, int *rec_sizes, int nr_recs) +{ + unsigned long long commit = 0; + int offset = TRACE_DAT_RECORD_HEADER_SIZE; + int i; + char *page; + + page = calloc(1, trace_dat_page_size); + if (!page) + return -ENOMEM; + + for (i = 0; i < nr_recs; i++) { + memcpy(page + offset, records[i], rec_sizes[i]); + offset += rec_sizes[i]; + commit += rec_sizes[i]; + } + + memcpy(page, &base_ts, sizeof(base_ts)); + memcpy(page + sizeof(base_ts), &commit, sizeof(commit)); + + if (!fwrite(page, 1, trace_dat_page_size, fp)) { + free(page); + return -EIO; + } + free(page); + + return 0; +} + +/* Write all trace data for a single CPU as trace.dat flyrecord pages */ +static int trace_dat__write_cpu_dat(FILE *fp, int cpu, unsigned long long *file_offset_out) +{ + struct cpu_events *cpu_events = &trace_cpu_data[cpu]; + unsigned long long base_ts; + unsigned long long file_offset; + char **page_records = NULL; + int *page_rec_sizes = NULL; + int page_cap = 0; + int nr_page_recs = 0; + int page_size_used = 0; + int ret = 0; + int i, j; + + file_offset = ftell(fp); + *file_offset_out = file_offset; + + if (cpu_events->count == 0) { + char *empty_page = calloc(1, trace_dat_page_size); + + if (!empty_page) + return -ENOMEM; + if (!fwrite(empty_page, 1, trace_dat_page_size, fp)) { + free(empty_page); + return -EIO; + } + free(empty_page); + return 0; + } + + base_ts = cpu_events->events[0].ts; + + for (i = 0; i < cpu_events->count; i++) { + struct cpu_event *event = &cpu_events->events[i]; + unsigned long long time_delta = event->ts - base_ts; + unsigned int data_len = event->raw_size; + unsigned int words; + unsigned int type_len; + unsigned int payload_offset; + unsigned int hdr_word; + char *extend = NULL; + int extend_size = 0; + char *data_rec; + int data_rec_size; + int needed_size; + + /* Emit TIME_EXTEND when delta does not fit in 27 bits */ + if (time_delta > TRACE_DAT_RECORD_TIME_MASK) { + unsigned int extend_hdr; + unsigned int delta_upper; + + extend_size = TRACE_DAT_RECORD_TIME_EXTEND_SIZE; + extend = calloc(1, extend_size); + if (!extend) + return -ENOMEM; + + extend_hdr = + ((time_delta & TRACE_DAT_RECORD_TIME_MASK) << + TRACE_DAT_RECORD_TIME_SHIFT) | + TRACE_DAT_RECORD_TYPE_TIME_EXTEND; + delta_upper = time_delta >> TRACE_DAT_RECORD_TIME_SHIFT; + + memcpy(extend, &extend_hdr, TRACE_DAT_WORD_SIZE); + memcpy(extend + TRACE_DAT_WORD_SIZE, &delta_upper, + TRACE_DAT_WORD_SIZE); + + time_delta = 0; + } + words = (data_len + TRACE_DAT_WORD_ALIGN_MASK) / TRACE_DAT_WORD_SIZE; + /* + * Encode event length per ftrace ring buffer format: + * - Small events (≤28 words): type_len = word count + * - Large events (≥29 words): type_len = 0, followed by byte length + */ + if (words <= TRACE_DAT_RECORD_TYPE_LEN_MAX) { + type_len = words; + payload_offset = TRACE_DAT_WORD_SIZE; + data_rec_size = TRACE_DAT_WORD_SIZE + data_len; + } else { + type_len = 0; + payload_offset = 2 * TRACE_DAT_WORD_SIZE; + data_rec_size = 2 * TRACE_DAT_WORD_SIZE + data_len; + } + + if (data_rec_size % TRACE_DAT_WORD_SIZE) + data_rec_size += TRACE_DAT_WORD_SIZE - + (data_rec_size % TRACE_DAT_WORD_SIZE); + + /* Reject records that cannot fit in a single page payload */ + if (data_rec_size > trace_dat_page_size - TRACE_DAT_RECORD_HEADER_SIZE) { + free(extend); + ret = -E2BIG; + goto out_free; + } + + needed_size = extend_size + data_rec_size; + + /* Check page fit BEFORE allocating data record */ + if (page_size_used + needed_size > + trace_dat_page_size - TRACE_DAT_RECORD_HEADER_SIZE) { + ret = trace_dat__write_page(fp, base_ts, + page_records, page_rec_sizes, + nr_page_recs); + + for (j = 0; j < nr_page_recs; j++) + free(page_records[j]); + + nr_page_recs = 0; + page_size_used = 0; + base_ts = event->ts; + + if (ret < 0) { + free(extend); + goto out_free; + } + + /* + * New page base matches this event timestamp, so no + * TIME_EXTEND is needed anymore. + */ + free(extend); + extend = NULL; + extend_size = 0; + time_delta = 0; + } + + hdr_word = (time_delta << TRACE_DAT_RECORD_TIME_SHIFT) | type_len; + + data_rec = calloc(1, data_rec_size); + if (!data_rec) { + free(extend); + ret = -ENOMEM; + goto out_free; + } + + memcpy(data_rec, &hdr_word, TRACE_DAT_WORD_SIZE); + + /* Large events: write actual byte length after header */ + if (type_len == 0) + memcpy(data_rec + TRACE_DAT_WORD_SIZE, &data_len, TRACE_DAT_WORD_SIZE); + + memcpy(data_rec + payload_offset, event->raw, data_len); + + if (nr_page_recs + (extend ? 1 : 0) >= page_cap) { + char **new_records; + int *new_sizes; + int needed = nr_page_recs + (extend ? 1 : 0) + 1; + int new_capacity = page_cap ? page_cap * 2 : + INITIAL_PAGE_RECORD_CAPACITY; + + while (new_capacity < needed) + new_capacity *= 2; + + new_records = malloc(new_capacity * sizeof(*new_records)); + new_sizes = malloc(new_capacity * sizeof(*new_sizes)); + if (!new_records || !new_sizes) { + free(new_records); + free(new_sizes); + free(extend); + free(data_rec); + ret = -ENOMEM; + goto out_free; + } + + if (nr_page_recs > 0) { + memcpy(new_records, page_records, + nr_page_recs * sizeof(*new_records)); + memcpy(new_sizes, page_rec_sizes, + nr_page_recs * sizeof(*new_sizes)); + } + + free(page_records); + free(page_rec_sizes); + + page_records = new_records; + page_rec_sizes = new_sizes; + page_cap = new_capacity; + } + + if (extend) { + page_records[nr_page_recs] = extend; + page_rec_sizes[nr_page_recs] = extend_size; + nr_page_recs++; + page_size_used += extend_size; + } + + page_records[nr_page_recs] = data_rec; + page_rec_sizes[nr_page_recs] = data_rec_size; + nr_page_recs++; + page_size_used += data_rec_size; + } + + if (nr_page_recs > 0) { + ret = trace_dat__write_page(fp, base_ts, + page_records, page_rec_sizes, nr_page_recs); + } +out_free: + for (j = 0; j < nr_page_recs; j++) + free(page_records[j]); + free(page_records); + free(page_rec_sizes); + return ret; +} + +/* Write the strings section containing section name lookup table */ +int trace_dat__write_strings_section(void) +{ + unsigned short section_id = TRACE_DAT_SECTION_STRINGS; + unsigned short flags = 0; + unsigned long long section_size = 0; + static const char * const section_names[] = { + "headers", /* offset 0 - strid for section 16 */ + "ftrace event formats", /* offset 8 - strid for section 17 */ + "events format", /* offset 29 - strid for section 18 */ + "kallsyms", /* offset 43 - strid for section 19 */ + "cmdlines", /* offset 52 - strid for section 21 */ + "strings", /* offset 61 - strid for section 15 */ + "options", /* offset 69 - strid for options 1 */ + "options", /* offset 77 - strid for options 2 */ + "buffer-flyrecord", /* offset 85 - strid for flyrecord 3 */ + NULL + }; + + /* string_id points to "strings" string itself */ + unsigned int string_id = STRID_STRINGS; + int i; + + if (!trace_dat_fp) + return -EBADF; + + for (i = 0; section_names[i] != NULL; i++) + section_size += strlen(section_names[i]) + 1; + + /* write section header */ + if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp) || + !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp)) + return -EIO; + + /* write strings */ + for (i = 0; section_names[i] != NULL; i++) + if (!fwrite(section_names[i], 1, strlen(section_names[i]) + 1, trace_dat_fp)) + return -EIO; + return 0; +} + +/* Writes options section containing CPUCOUNT, TRACECLOCK, EVENT_FORMAT, HEADER_INFO, + * FTRACE_EVENTS, KALLSYMS, CMDLINES options, ending with DONE option pointing to next section. + */ +int trace_dat__write_options_section1(void) +{ + unsigned short section_id = TRACE_DAT_SECTION_OPTIONS; + unsigned short flags = 0; + unsigned int string_id = STRID_OPTIONS_1; + unsigned long long section_size = 0; + long section_size_pos; + long payload_start; + unsigned long long section_start; + unsigned short opt_id; + unsigned int opt_size; + char clock_buf[CLOCK_BUFFER_SIZE]; + FILE *clock_file; + size_t bytes_read; + char *path; + unsigned long long next_offset; + long end_pos; + + if (!trace_dat_fp) + return -EBADF; + + /* fill options_offset in initial format */ + section_start = ftell(trace_dat_fp); + + if (fseek(trace_dat_fp, trace_dat_options_offset, SEEK_SET) < 0 || + !fwrite(§ion_start, sizeof(unsigned long long), 1, trace_dat_fp) || + fseek(trace_dat_fp, 0, SEEK_END) < 0) + return -EIO; + + /* write section header */ + if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp)) + return -EIO; + section_size_pos = ftell(trace_dat_fp); + if (!fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp)) + return -EIO; + + payload_start = ftell(trace_dat_fp); + + /* CPUCOUNT option */ + opt_id = TRACE_DAT_OPTION_CPUCOUNT; + opt_size = sizeof(unsigned int); + + if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) || + !fwrite(&trace_dat_nr_cpus, sizeof(unsigned int), 1, trace_dat_fp)) + return -EIO; + + /* TRACECLOCK option */ + opt_id = TRACE_DAT_OPTION_TRACECLOCK; + + path = get_tracing_file("trace_clock"); + if (path) { + clock_file = fopen(path, "r"); + put_tracing_file(path); + } else { + clock_file = NULL; + } + if (clock_file) { + bytes_read = fread(clock_buf, 1, sizeof(clock_buf) - 1, clock_file); + fclose(clock_file); + clock_buf[bytes_read] = '\0'; + } else { + strcpy(clock_buf, "local\n"); + bytes_read = strlen(clock_buf); + } + opt_size = bytes_read + 1; + if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) || + !fwrite(clock_buf, 1, opt_size, trace_dat_fp)) + return -EIO; + + /* EVENT option */ + opt_id = TRACE_DAT_OPTION_EVENT; + opt_size = sizeof(unsigned long long); + + if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) || + !fwrite(&trace_dat_events_format_offset, sizeof(unsigned long long), + 1, trace_dat_fp)) + return -EIO; + + /* HEADER option */ + opt_id = TRACE_DAT_OPTION_HEADER; + opt_size = sizeof(unsigned long long); + + if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) || + !fwrite(&trace_dat_header_info_offset, sizeof(unsigned long long), + 1, trace_dat_fp)) + return -EIO; + + /* FTRACE option */ + opt_id = TRACE_DAT_OPTION_FTRACE; + opt_size = sizeof(unsigned long long); + + if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) || + !fwrite(&trace_dat_ftrace_format_offset, sizeof(unsigned long long), + 1, trace_dat_fp)) + return -EIO; + + /* KALLSYMS option */ + opt_id = TRACE_DAT_OPTION_KALLSYMS; + opt_size = sizeof(unsigned long long); + + if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) || + !fwrite(&trace_dat_kallsyms_offset, sizeof(unsigned long long), + 1, trace_dat_fp)) + return -EIO; + + /* CMDLINE option */ + opt_id = TRACE_DAT_OPTION_CMDLINE; + opt_size = sizeof(unsigned long long); + + if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) || + !fwrite(&trace_dat_cmdline_offset, sizeof(unsigned long long), + 1, trace_dat_fp)) + return -EIO; + + /* DONE option id - next_options_offset filled later */ + opt_id = TRACE_DAT_OPTION_DONE; + opt_size = sizeof(unsigned long long); + next_offset = 0; /* placeholder */ + + trace_dat_next_options_offset = ftell(trace_dat_fp); + if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) || + !fwrite(&next_offset, sizeof(unsigned long long), 1, trace_dat_fp)) + return -EIO; + + /* fill section size */ + end_pos = ftell(trace_dat_fp); + + section_size = end_pos - payload_start; + if (fseek(trace_dat_fp, section_size_pos, SEEK_SET) < 0 || + !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp) || + fseek(trace_dat_fp, end_pos, SEEK_SET) < 0) + return -EIO; + + return 0; + +} + +/* Writes options section containing BUFFER option with flyrecord section + * (flyrecord section offset, clock type, page size, CPU count, + * per-CPU offsets/sizes) and DONE option. + */ +int trace_dat__write_options_section2(void) +{ + unsigned short section_id = TRACE_DAT_SECTION_OPTIONS; + unsigned short flags = 0; + unsigned int string_id = STRID_OPTIONS_2; + unsigned long long section_size = 0; + long section_size_pos; + long payload_start; + int cpu; + unsigned short opt_id = TRACE_DAT_OPTION_BUFFER; + unsigned int opt_size = 0; + long opt_size_pos; + unsigned long long data_offset = 0; + unsigned int page_size = (unsigned int)trace_dat_page_size; + const char *clock = "local"; + unsigned long long next; + long end_pos; + unsigned long long cpu_offset; + unsigned long long cpu_size; + unsigned short done_id; + unsigned int done_size; + + if (!trace_dat_fp) + return -EINVAL; + + /* fill done1 next offset - points to this section */ + next = ftell(trace_dat_fp); + + if (fseek(trace_dat_fp, trace_dat_next_options_offset + 2 + 4, SEEK_SET) < 0 || + !fwrite(&next, sizeof(unsigned long long), 1, trace_dat_fp) || + fseek(trace_dat_fp, 0, SEEK_END) < 0) + return -EIO; + + /* write section header */ + if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp)) + return -EIO; + section_size_pos = ftell(trace_dat_fp); + if (!fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp)) + return -EIO; + + payload_start = ftell(trace_dat_fp); + + /* BUFFER option */ + if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp)) + return -EIO; + opt_size_pos = ftell(trace_dat_fp); + if (!fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp)) + return -EIO; + opt_payload_start = ftell(trace_dat_fp); + + /* data_offset placeholder */ + if (!fwrite(&data_offset, sizeof(unsigned long long), 1, trace_dat_fp) || + !fwrite("\0", 1, 1, trace_dat_fp) || + !fwrite(clock, 1, strlen(clock) + 1, trace_dat_fp) || + !fwrite(&page_size, sizeof(unsigned int), 1, trace_dat_fp) || + !fwrite(&trace_dat_nr_cpus, sizeof(unsigned int), 1, trace_dat_fp)) + return -EIO; + + /* per cpu: cpu_id + offset placeholder + size */ + for (cpu = 0; cpu < trace_dat_nr_cpus; cpu++) { + cpu_offset = 0; /* filled in write_flyrecord */ + cpu_size = 0; /* filled in write_flyrecord */ + + if (!fwrite(&cpu, sizeof(unsigned int), 1, trace_dat_fp)) + return -EIO; + buffer_opt_cpu_offsets_pos[cpu] = ftell(trace_dat_fp); + if (!fwrite(&cpu_offset, sizeof(unsigned long long), 1, trace_dat_fp) || + !fwrite(&cpu_size, sizeof(unsigned long long), 1, trace_dat_fp)) + return -EIO; + } + + /* fill opt_size */ + end_pos = ftell(trace_dat_fp); + + opt_size = end_pos - opt_payload_start; + fseek(trace_dat_fp, opt_size_pos, SEEK_SET); + if (!fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp)) + return -EIO; + fseek(trace_dat_fp, end_pos, SEEK_SET); + + /* DONE id=0 */ + done_id = TRACE_DAT_OPTION_DONE; + done_size = sizeof(unsigned long long); + /* No additional options sections follow this one */ + next = 0; + + if (!fwrite(&done_id, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&done_size, sizeof(unsigned int), 1, trace_dat_fp) || + !fwrite(&next, sizeof(unsigned long long), 1, trace_dat_fp)) + return -EIO; + + /* fill section size */ + end_pos = ftell(trace_dat_fp); + + section_size = end_pos - payload_start; + fseek(trace_dat_fp, section_size_pos, SEEK_SET); + if (!fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp)) + return -EIO; + fseek(trace_dat_fp, end_pos, SEEK_SET); + + return 0; + +} + +int trace_dat__write_flyrecord_section(void) +{ + unsigned short section_id = TRACE_DAT_SECTION_FLYRECORD; + unsigned short flags = 0; + unsigned int string_id = STRID_BUFFER_FLYRECORD; + unsigned long long section_size = 0; + long section_size_pos; + long flyrecord_start; + long after_header; + long padding_needed; + unsigned long long *cpu_offsets; + unsigned long long *cpu_sizes; + int cpu; + int ret = 0; + char *pad; + unsigned long long start; + long end_pos; + + if (!trace_dat_fp) + return -EINVAL; + + cpu_offsets = calloc(trace_dat_nr_cpus, sizeof(unsigned long long)); + cpu_sizes = calloc(trace_dat_nr_cpus, sizeof(unsigned long long)); + if (!cpu_offsets || !cpu_sizes) { + ret = -ENOMEM; + goto cleanup; + } + flyrecord_start = ftell(trace_dat_fp); + if (flyrecord_start < 0) { + ret = -EIO; + goto cleanup; + } + + /* section header */ + if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp)) { + ret = -EIO; + goto cleanup; + } + section_size_pos = ftell(trace_dat_fp); + if (!fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp)) { + ret = -EIO; + goto cleanup; + } + + /* Align to page boundary */ + after_header = ftell(trace_dat_fp); + padding_needed = (trace_dat_page_size - + (after_header % trace_dat_page_size)) % trace_dat_page_size; + + if (padding_needed > 0) { + pad = calloc(1, padding_needed); + if (!pad) { + ret = -ENOMEM; + goto cleanup; + } + + if (!fwrite(pad, 1, padding_needed, trace_dat_fp)) { + free(pad); + ret = -EIO; + goto cleanup; + } + free(pad); + } + + /* write per-cpu trace data */ + for (cpu = 0; cpu < trace_dat_nr_cpus; cpu++) { + start = ftell(trace_dat_fp); + + ret = trace_dat__write_cpu_dat(trace_dat_fp, cpu, &cpu_offsets[cpu]); + + if (ret < 0) { + pr_err("Failed to write CPU %d data\n", cpu); + goto cleanup; + } + cpu_sizes[cpu] = ftell(trace_dat_fp) - start; + } + + /* fill section size */ + end_pos = ftell(trace_dat_fp); + + section_size = end_pos - after_header; + if (fseek(trace_dat_fp, section_size_pos, SEEK_SET) < 0 || + !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp)) { + ret = -EIO; + goto cleanup; + } + if (fseek(trace_dat_fp, end_pos, SEEK_SET) < 0) { + ret = -EIO; + goto cleanup; + } + + /* fill cpu offsets and sizes in BUFFER option */ + for (cpu = 0; cpu < trace_dat_nr_cpus; cpu++) { + if (fseek(trace_dat_fp, buffer_opt_cpu_offsets_pos[cpu], SEEK_SET) < 0 || + !fwrite(&cpu_offsets[cpu], sizeof(unsigned long long), 1, trace_dat_fp) || + !fwrite(&cpu_sizes[cpu], sizeof(unsigned long long), 1, trace_dat_fp)) { + ret = -EIO; + goto cleanup; + } + } + + /* fill data offset in buffer option */ + if (fseek(trace_dat_fp, opt_payload_start, SEEK_SET) < 0 || + !fwrite(&flyrecord_start, sizeof(unsigned long long), 1, trace_dat_fp)) { + ret = -EIO; + goto cleanup; + } + + if (fseek(trace_dat_fp, 0, SEEK_END) < 0) { + ret = -EIO; + goto cleanup; + } + + +cleanup: + free(cpu_offsets); + free(cpu_sizes); + return ret; +} + +/* Free all per-CPU event buffers */ +void trace_dat__free_cpu_buffers(void) +{ + int cpu; + + if (!trace_cpu_data) + return; + + for (cpu = 0; cpu < trace_dat_nr_cpus; cpu++) { + int i; + + for (i = 0; i < trace_cpu_data[cpu].count; i++) + free(trace_cpu_data[cpu].events[i].raw); + free(trace_cpu_data[cpu].events); + } + free(trace_cpu_data); + trace_cpu_data = NULL; + free(buffer_opt_cpu_offsets_pos); + buffer_opt_cpu_offsets_pos = NULL; + trace_dat_nr_cpus = 0; +} diff --git a/tools/perf/util/trace-dat.h b/tools/perf/util/trace-dat.h new file mode 100644 index 000000000000..9aec37b708d4 --- /dev/null +++ b/tools/perf/util/trace-dat.h @@ -0,0 +1,83 @@ +/* SPDX-License-Identifier: GPL-2.0-or-later */ +/* + * Copyright 2026, IBM Corporation + * Author: Tanushree Shah <tshah@linux.ibm.com> + */ + +#ifndef __PERF_TRACE_DAT_H +#define __PERF_TRACE_DAT_H + +#include <stdio.h> + +/* trace.dat file format version */ +#define TRACE_DAT_VERSION '7' + +/* + * Section IDs for trace.dat format + */ +#define TRACE_DAT_SECTION_OPTIONS 0 +#define TRACE_DAT_SECTION_FLYRECORD 3 +#define TRACE_DAT_SECTION_STRINGS 15 +#define TRACE_DAT_SECTION_HEADER 16 +#define TRACE_DAT_SECTION_FTRACE 17 +#define TRACE_DAT_SECTION_EVENTS 18 +#define TRACE_DAT_SECTION_KALLSYMS 19 +#define TRACE_DAT_SECTION_CMDLINE 21 + +/* + * Option IDs for trace.dat options sections + */ +#define TRACE_DAT_OPTION_DONE 0 +#define TRACE_DAT_OPTION_BUFFER 3 +#define TRACE_DAT_OPTION_TRACECLOCK 4 +#define TRACE_DAT_OPTION_CPUCOUNT 8 +#define TRACE_DAT_OPTION_HEADER 16 +#define TRACE_DAT_OPTION_FTRACE 17 +#define TRACE_DAT_OPTION_EVENT 18 +#define TRACE_DAT_OPTION_KALLSYMS 19 +#define TRACE_DAT_OPTION_CMDLINE 21 + +/* + * String offsets in the strings section + * These point to null-terminated strings used as section names + */ +#define STRID_HEADERS 0 +#define STRID_FTRACE_FORMATS 8 +#define STRID_EVENT_FORMATS 29 +#define STRID_KALLSYMS 43 +#define STRID_CMDLINES 52 +#define STRID_STRINGS 61 +#define STRID_OPTIONS_1 69 +#define STRID_OPTIONS_2 77 +#define STRID_BUFFER_FLYRECORD 85 + +struct perf_session; + +extern FILE *trace_dat_fp; +extern int trace_dat_page_size; +extern int trace_dat_nr_cpus; +extern long trace_dat_options_offset; +extern long trace_dat_header_info_offset; +extern long trace_dat_events_format_offset; +extern long trace_dat_ftrace_format_offset; +extern long trace_dat_kallsyms_offset; +extern long trace_dat_cmdline_offset; +extern long trace_dat_next_options_offset; + +/* collect and manage per-cpu tracepoint event buffers */ +#define TRACE_DAT_RECORD_TYPE_TIME_EXTEND 30 +#define TRACE_DAT_RECORD_TIME_EXTEND_SIZE 8 +#define TRACE_DAT_RECORD_TYPE_LEN_MAX 28 + +int trace_dat__init_cpu_buffers(int nr_cpus); +int trace_dat__collect_cpu_event(int cpu, unsigned long long ts, + void *raw, unsigned int raw_size); +void trace_dat__free_cpu_buffers(void); + +/* write trace.dat file sections */ +int trace_dat__write_options_section1(void); +int trace_dat__write_options_section2(void); +int trace_dat__write_flyrecord_section(void); +int trace_dat__write_strings_section(void); + +#endif /* __PERF_TRACE_DAT_H */ -- 2.47.1 ^ permalink raw reply related [flat|nested] 14+ messages in thread
* Re: [RFC PATCH v3 1/5] perf/trace-dat: Add trace.dat export infrastructure 2026-08-03 14:59 ` [RFC PATCH v3 1/5] perf/trace-dat: Add trace.dat export infrastructure Tanushree Shah @ 2026-08-03 15:14 ` sashiko-bot 0 siblings, 0 replies; 14+ messages in thread From: sashiko-bot @ 2026-08-03 15:14 UTC (permalink / raw) To: Tanushree Shah; +Cc: linux-perf-users Thank you for your contribution! Sashiko AI review found 5 potential issue(s) to consider: - [High] The `TIME_EXTEND` upper bits are calculated using an incorrect bitwise right shift of 5 instead of 27, resulting in corrupted extended timestamps. - [High] The ftrace ring buffer event length encoding for large events incorrectly writes the unaligned payload length, omitting necessary alignment padding. - [High] Using 32-bit `long` variables and `ftell()`/`fseek()` for file offsets truncates trace files larger than 2GB on 32-bit systems. - [High] Trace data export file headers and ftrace metadata are written in host endianness instead of the target trace's endianness, breaking cross-platform analysis. - [Medium] Lack of validation for `trace_dat_page_size` can result in a fatal integer division by zero exception when calculating page padding. -- --- Patch [1]: [RFC PATCH v3 1/5] perf/trace-dat: Add trace.dat export infrastructure --- commit 34a2391aa9ca8393dfbdda4a477f39710a409324 Author: Tanushree Shah <tshah@linux.ibm.com> perf/trace-dat: Add trace.dat export infrastructure This commit adds new utility files to implement infrastructure for exporting perf.data tracepoints to the trace.dat format. It provides functions for managing per-cpu event buffers and constructing ftrace ring buffer pages. > diff --git a/tools/perf/util/trace-dat.c b/tools/perf/util/trace-dat.c > new file mode 100644 > index 0000000000000..55a7bd982c6f8 > --- /dev/null > +++ b/tools/perf/util/trace-dat.c [ ... ] > +static int trace_dat__write_page(FILE *fp, unsigned long long base_ts, > + char **records, int *rec_sizes, int nr_recs) > +{ > + unsigned long long commit = 0; > + int offset = TRACE_DAT_RECORD_HEADER_SIZE; > + int i; > + char *page; > + > + page = calloc(1, trace_dat_page_size); > + if (!page) > + return -ENOMEM; > + > + for (i = 0; i < nr_recs; i++) { > + memcpy(page + offset, records[i], rec_sizes[i]); > + offset += rec_sizes[i]; > + commit += rec_sizes[i]; > + } > + > + memcpy(page, &base_ts, sizeof(base_ts)); > + memcpy(page + sizeof(base_ts), &commit, sizeof(commit)); [Severity: High] Does this write the ftrace page header using native host endianness? If a trace was recorded on a different architecture, copying the integers directly to the output stream might mix host-endian metadata with target-endian payloads. Should this metadata be byte-swapped to match the target architecture to ensure cross-platform trace analysis works correctly? > + > + if (!fwrite(page, 1, trace_dat_page_size, fp)) { > + free(page); > + return -EIO; > + } > + free(page); > + > + return 0; > +} [ ... ] > +static int trace_dat__write_cpu_dat(FILE *fp, int cpu, unsigned long long *file_offset_out) > +{ [ ... ] > + /* Emit TIME_EXTEND when delta does not fit in 27 bits */ > + if (time_delta > TRACE_DAT_RECORD_TIME_MASK) { > + unsigned int extend_hdr; > + unsigned int delta_upper; > + > + extend_size = TRACE_DAT_RECORD_TIME_EXTEND_SIZE; > + extend = calloc(1, extend_size); > + if (!extend) > + return -ENOMEM; > + > + extend_hdr = > + ((time_delta & TRACE_DAT_RECORD_TIME_MASK) << > + TRACE_DAT_RECORD_TIME_SHIFT) | > + TRACE_DAT_RECORD_TYPE_TIME_EXTEND; > + delta_upper = time_delta >> TRACE_DAT_RECORD_TIME_SHIFT; [Severity: High] Does this shift the upper bits of the time delta by the correct amount? The code uses TRACE_DAT_RECORD_TIME_SHIFT, which is 5. This calculates the upper bits as time_delta >> 5 rather than >> 27. Would this copy the middle 22 bits instead of the upper bits, causing timestamp corruption for any events that follow a time extend? > + > + memcpy(extend, &extend_hdr, TRACE_DAT_WORD_SIZE); > + memcpy(extend + TRACE_DAT_WORD_SIZE, &delta_upper, > + TRACE_DAT_WORD_SIZE); > + > + time_delta = 0; > + } [ ... ] > + hdr_word = (time_delta << TRACE_DAT_RECORD_TIME_SHIFT) | type_len; > + > + data_rec = calloc(1, data_rec_size); > + if (!data_rec) { > + free(extend); > + ret = -ENOMEM; > + goto out_free; > + } > + > + memcpy(data_rec, &hdr_word, TRACE_DAT_WORD_SIZE); > + > + /* Large events: write actual byte length after header */ > + if (type_len == 0) > + memcpy(data_rec + TRACE_DAT_WORD_SIZE, &data_len, TRACE_DAT_WORD_SIZE); [Severity: High] For large events, is this writing the raw, unaligned data_len into the payload length field? The internal offset data_rec_size advances by the padded size, but the unaligned data_len is written here. If data_len is not a multiple of 4, a parser reading this file might compute an incorrect total length and fail to skip the alignment padding. Could this cause the reader to misread the next event boundary? > + > + memcpy(data_rec + payload_offset, event->raw, data_len); > + [ ... ] > +int trace_dat__write_flyrecord_section(void) > +{ [ ... ] > + section_size_pos = ftell(trace_dat_fp); > + if (!fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp)) { > + ret = -EIO; > + goto cleanup; > + } > + > + /* Align to page boundary */ > + after_header = ftell(trace_dat_fp); > + padding_needed = (trace_dat_page_size - > + (after_header % trace_dat_page_size)) % trace_dat_page_size; [Severity: Medium] Is trace_dat_page_size guaranteed to be non-zero here? If the caller fails to initialize this global variable or parses a configuration that sets it to zero, this modulo operation will trigger a fatal integer division by zero exception. > + > + if (padding_needed > 0) { > + pad = calloc(1, padding_needed); > + if (!pad) { > + ret = -ENOMEM; > + goto cleanup; > + } > + > + if (!fwrite(pad, 1, padding_needed, trace_dat_fp)) { > + free(pad); > + ret = -EIO; > + goto cleanup; > + } > + free(pad); > + } > + > + /* write per-cpu trace data */ [ ... ] > diff --git a/tools/perf/util/trace-dat.h b/tools/perf/util/trace-dat.h > new file mode 100644 > index 0000000000000..9aec37b708d41 > --- /dev/null > +++ b/tools/perf/util/trace-dat.h > @@ -0,0 +1,83 @@ [ ... ] > +extern FILE *trace_dat_fp; > +extern int trace_dat_page_size; > +extern int trace_dat_nr_cpus; > +extern long trace_dat_options_offset; > +extern long trace_dat_header_info_offset; > +extern long trace_dat_events_format_offset; > +extern long trace_dat_ftrace_format_offset; > +extern long trace_dat_kallsyms_offset; > +extern long trace_dat_cmdline_offset; > +extern long trace_dat_next_options_offset; [Severity: High] Are these offsets safe for trace files larger than 2GB on 32-bit systems? Using long for offsets and relying on ftell() limits the maximum supported file size to 2GB on 32-bit architectures. Since perf.data files can easily exceed this size, could this result in integer overflow and corrupted offsets? Should this use standard 64-bit interfaces like u64 and ftello()/fseeko() instead? > + > +/* collect and manage per-cpu tracepoint event buffers */ > +#define TRACE_DAT_RECORD_TYPE_TIME_EXTEND 30 > +#define TRACE_DAT_RECORD_TIME_EXTEND_SIZE 8 > +#define TRACE_DAT_RECORD_TYPE_LEN_MAX 28 [ ... ] -- Sashiko AI review · https://sashiko.dev/#/patchset/20260803145958.299956-1-tshah@linux.ibm.com?part=1 ^ permalink raw reply [flat|nested] 14+ messages in thread
* [RFC PATCH v3 2/5] perf/trace-event: Write trace.dat metadata sections during parsing 2026-08-03 14:59 [RFC PATCH v3 0/5] Add perf.data tracepoint events to trace.dat conversion Tanushree Shah 2026-08-03 14:59 ` [RFC PATCH v3 1/5] perf/trace-dat: Add trace.dat export infrastructure Tanushree Shah @ 2026-08-03 14:59 ` Tanushree Shah 2026-08-03 15:14 ` sashiko-bot 2026-08-03 14:59 ` [RFC PATCH v3 3/5] perf data-convert: Add perf.data to trace.dat conversion backend Tanushree Shah ` (3 subsequent siblings) 5 siblings, 1 reply; 14+ messages in thread From: Tanushree Shah @ 2026-08-03 14:59 UTC (permalink / raw) To: acme, jolsa, adrian.hunter, vmolnaro, mpetlan, tmricht, maddy, irogers, namhyung Cc: linux-perf-users, linuxppc-dev, atrajeev, hbathini, Tejas.Manhas1, Tanushree.Shah, rostedt, Tanushree Shah Perf already captures the tracing metadata as a part of data section in perf.data When trace_dat_fp is set, write trace.dat compatible metadata sections using the perf provided raw buffers. Sections written: - Initial format header (magic, version, endian, long_size, page_size, compression, options_offset placeholder) - Section 16: HEADER INFO (header_page + header_event) - Section 17: FTRACE EVENT FORMATS - Section 18: EVENT FORMATS (per system/event format files) - Section 19: KALLSYMS - Section 21: CMDLINES - Section 15: STRINGS (written last after all sections) Introduce to_file_u16/u32/u64 helpers wrapping tep_read_number() to write all trace.dat metadata in the recorded machine's byte order, ensuring correct cross-platform compatibility. Signed-off-by: Tanushree Shah <tshah@linux.ibm.com> --- tools/perf/util/trace-dat.c | 212 ++++++++++++-------- tools/perf/util/trace-dat.h | 40 +++- tools/perf/util/trace-event-read.c | 307 ++++++++++++++++++++++++++++- 3 files changed, 465 insertions(+), 94 deletions(-) diff --git a/tools/perf/util/trace-dat.c b/tools/perf/util/trace-dat.c index 55a7bd982c6f..f71e03716e27 100644 --- a/tools/perf/util/trace-dat.c +++ b/tools/perf/util/trace-dat.c @@ -43,6 +43,7 @@ /* Buffer size for reading trace_clock string from debugfs/tracefs */ #define CLOCK_BUFFER_SIZE 256 +bool trace_dat_write_failed; FILE *trace_dat_fp; int trace_dat_page_size; int trace_dat_nr_cpus; @@ -143,10 +144,12 @@ int trace_dat__collect_cpu_event(int cpu, unsigned long long ts, } /* Write a single page of trace records */ -static int trace_dat__write_page(FILE *fp, unsigned long long base_ts, +static int trace_dat__write_page(FILE *fp, struct tep_handle *pevent, unsigned long long base_ts, char **records, int *rec_sizes, int nr_recs) { unsigned long long commit = 0; + unsigned long long ts_out; + unsigned long long commit_out; int offset = TRACE_DAT_RECORD_HEADER_SIZE; int i; char *page; @@ -161,8 +164,12 @@ static int trace_dat__write_page(FILE *fp, unsigned long long base_ts, commit += rec_sizes[i]; } - memcpy(page, &base_ts, sizeof(base_ts)); - memcpy(page + sizeof(base_ts), &commit, sizeof(commit)); + /* Byte-swap page header for cross-arch compatibility */ + ts_out = to_file_u64(pevent, base_ts); + commit_out = to_file_u64(pevent, commit); + + memcpy(page, &ts_out, sizeof(ts_out)); + memcpy(page + sizeof(ts_out), &commit_out, sizeof(commit_out)); if (!fwrite(page, 1, trace_dat_page_size, fp)) { free(page); @@ -174,7 +181,8 @@ static int trace_dat__write_page(FILE *fp, unsigned long long base_ts, } /* Write all trace data for a single CPU as trace.dat flyrecord pages */ -static int trace_dat__write_cpu_dat(FILE *fp, int cpu, unsigned long long *file_offset_out) +static int trace_dat__write_cpu_dat(FILE *fp, struct tep_handle *pevent, + int cpu, unsigned long long *file_offset_out) { struct cpu_events *cpu_events = &trace_cpu_data[cpu]; unsigned long long base_ts; @@ -229,11 +237,18 @@ static int trace_dat__write_cpu_dat(FILE *fp, int cpu, unsigned long long *file_ if (!extend) return -ENOMEM; - extend_hdr = - ((time_delta & TRACE_DAT_RECORD_TIME_MASK) << - TRACE_DAT_RECORD_TIME_SHIFT) | - TRACE_DAT_RECORD_TYPE_TIME_EXTEND; - delta_upper = time_delta >> TRACE_DAT_RECORD_TIME_SHIFT; + if (tep_is_file_bigendian(pevent)) { + extend_hdr = (time_delta & TRACE_DAT_RECORD_TIME_MASK) | + (TRACE_DAT_RECORD_TYPE_TIME_EXTEND << 27); + delta_upper = time_delta >> 27; + } else { + extend_hdr = ((time_delta & TRACE_DAT_RECORD_TIME_MASK) << + TRACE_DAT_RECORD_TIME_SHIFT) | + TRACE_DAT_RECORD_TYPE_TIME_EXTEND; + delta_upper = time_delta >> TRACE_DAT_RECORD_TIME_SHIFT; + } + extend_hdr = to_file_u32(pevent, extend_hdr); /* still needed */ + delta_upper = to_file_u32(pevent, delta_upper); /* still needed */ memcpy(extend, &extend_hdr, TRACE_DAT_WORD_SIZE); memcpy(extend + TRACE_DAT_WORD_SIZE, &delta_upper, @@ -273,7 +288,7 @@ static int trace_dat__write_cpu_dat(FILE *fp, int cpu, unsigned long long *file_ /* Check page fit BEFORE allocating data record */ if (page_size_used + needed_size > trace_dat_page_size - TRACE_DAT_RECORD_HEADER_SIZE) { - ret = trace_dat__write_page(fp, base_ts, + ret = trace_dat__write_page(fp, pevent, base_ts, page_records, page_rec_sizes, nr_page_recs); @@ -299,7 +314,13 @@ static int trace_dat__write_cpu_dat(FILE *fp, int cpu, unsigned long long *file_ time_delta = 0; } - hdr_word = (time_delta << TRACE_DAT_RECORD_TIME_SHIFT) | type_len; + if (tep_is_file_bigendian(pevent)) + hdr_word = ((unsigned int)time_delta & TRACE_DAT_RECORD_TIME_MASK) | + (type_len << 27); + else + hdr_word = ((unsigned int)time_delta << TRACE_DAT_RECORD_TIME_SHIFT) | + type_len; + hdr_word = to_file_u32(pevent, hdr_word); data_rec = calloc(1, data_rec_size); if (!data_rec) { @@ -311,8 +332,11 @@ static int trace_dat__write_cpu_dat(FILE *fp, int cpu, unsigned long long *file_ memcpy(data_rec, &hdr_word, TRACE_DAT_WORD_SIZE); /* Large events: write actual byte length after header */ - if (type_len == 0) - memcpy(data_rec + TRACE_DAT_WORD_SIZE, &data_len, TRACE_DAT_WORD_SIZE); + if (type_len == 0) { + unsigned int data_len_out = to_file_u32(pevent, data_len); + + memcpy(data_rec + TRACE_DAT_WORD_SIZE, &data_len_out, TRACE_DAT_WORD_SIZE); + } memcpy(data_rec + payload_offset, event->raw, data_len); @@ -366,7 +390,7 @@ static int trace_dat__write_cpu_dat(FILE *fp, int cpu, unsigned long long *file_ } if (nr_page_recs > 0) { - ret = trace_dat__write_page(fp, base_ts, + ret = trace_dat__write_page(fp, pevent, base_ts, page_records, page_rec_sizes, nr_page_recs); } out_free: @@ -378,11 +402,11 @@ static int trace_dat__write_cpu_dat(FILE *fp, int cpu, unsigned long long *file_ } /* Write the strings section containing section name lookup table */ -int trace_dat__write_strings_section(void) +int trace_dat__write_strings_section(struct tep_handle *pevent) { - unsigned short section_id = TRACE_DAT_SECTION_STRINGS; - unsigned short flags = 0; unsigned long long section_size = 0; + unsigned short section_id = to_file_u16(pevent, TRACE_DAT_SECTION_STRINGS); + unsigned short flags = to_file_u16(pevent, 0); static const char * const section_names[] = { "headers", /* offset 0 - strid for section 16 */ "ftrace event formats", /* offset 8 - strid for section 17 */ @@ -397,7 +421,7 @@ int trace_dat__write_strings_section(void) }; /* string_id points to "strings" string itself */ - unsigned int string_id = STRID_STRINGS; + unsigned int string_id = to_file_u32(pevent, STRID_STRINGS); int i; if (!trace_dat_fp) @@ -405,47 +429,59 @@ int trace_dat__write_strings_section(void) for (i = 0; section_names[i] != NULL; i++) section_size += strlen(section_names[i]) + 1; + section_size = to_file_u64(pevent, section_size); /* write section header */ if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) || !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) || !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp) || - !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp)) - return -EIO; + !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp)) { + pr_warning("Failed to write strings section\n"); + trace_dat_write_failed = true; + } /* write strings */ for (i = 0; section_names[i] != NULL; i++) - if (!fwrite(section_names[i], 1, strlen(section_names[i]) + 1, trace_dat_fp)) - return -EIO; + if (!fwrite(section_names[i], 1, strlen(section_names[i]) + 1, trace_dat_fp)) { + pr_warning("Failed to write strings section\n"); + trace_dat_write_failed = true; + } return 0; } /* Writes options section containing CPUCOUNT, TRACECLOCK, EVENT_FORMAT, HEADER_INFO, * FTRACE_EVENTS, KALLSYMS, CMDLINES options, ending with DONE option pointing to next section. */ -int trace_dat__write_options_section1(void) +int trace_dat__write_options_section1(struct tep_handle *pevent) { - unsigned short section_id = TRACE_DAT_SECTION_OPTIONS; - unsigned short flags = 0; - unsigned int string_id = STRID_OPTIONS_1; + unsigned short section_id = to_file_u16(pevent, TRACE_DAT_SECTION_OPTIONS); + unsigned short flags = to_file_u16(pevent, 0); + unsigned int string_id = to_file_u32(pevent, STRID_OPTIONS_1); unsigned long long section_size = 0; long section_size_pos; long payload_start; unsigned long long section_start; unsigned short opt_id; unsigned int opt_size; + unsigned int raw_opt_size; + unsigned int nr_cpus; char clock_buf[CLOCK_BUFFER_SIZE]; FILE *clock_file; size_t bytes_read; char *path; unsigned long long next_offset; + unsigned long long events_off; + unsigned long long header_off; + unsigned long long ftrace_off; + unsigned long long kallsyms_off; + unsigned long long cmdline_off; long end_pos; if (!trace_dat_fp) return -EBADF; /* fill options_offset in initial format */ - section_start = ftell(trace_dat_fp); + section_start = to_file_u64(pevent, ftell(trace_dat_fp)); if (fseek(trace_dat_fp, trace_dat_options_offset, SEEK_SET) < 0 || !fwrite(§ion_start, sizeof(unsigned long long), 1, trace_dat_fp) || @@ -464,16 +500,17 @@ int trace_dat__write_options_section1(void) payload_start = ftell(trace_dat_fp); /* CPUCOUNT option */ - opt_id = TRACE_DAT_OPTION_CPUCOUNT; - opt_size = sizeof(unsigned int); + opt_id = to_file_u16(pevent, TRACE_DAT_OPTION_CPUCOUNT); + opt_size = to_file_u32(pevent, sizeof(unsigned int)); + nr_cpus = to_file_u32(pevent, trace_dat_nr_cpus); if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) || !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) || - !fwrite(&trace_dat_nr_cpus, sizeof(unsigned int), 1, trace_dat_fp)) + !fwrite(&nr_cpus, sizeof(unsigned int), 1, trace_dat_fp)) return -EIO; /* TRACECLOCK option */ - opt_id = TRACE_DAT_OPTION_TRACECLOCK; + opt_id = to_file_u16(pevent, TRACE_DAT_OPTION_TRACECLOCK); path = get_tracing_file("trace_clock"); if (path) { @@ -490,66 +527,72 @@ int trace_dat__write_options_section1(void) strcpy(clock_buf, "local\n"); bytes_read = strlen(clock_buf); } - opt_size = bytes_read + 1; + raw_opt_size = bytes_read + 1; + opt_size = to_file_u32(pevent, bytes_read + 1); if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) || !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) || - !fwrite(clock_buf, 1, opt_size, trace_dat_fp)) + !fwrite(clock_buf, 1, raw_opt_size, trace_dat_fp)) return -EIO; /* EVENT option */ - opt_id = TRACE_DAT_OPTION_EVENT; - opt_size = sizeof(unsigned long long); + opt_id = to_file_u16(pevent, TRACE_DAT_OPTION_EVENT); + opt_size = to_file_u32(pevent, sizeof(unsigned long long)); + events_off = to_file_u64(pevent, trace_dat_events_format_offset); if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) || !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) || - !fwrite(&trace_dat_events_format_offset, sizeof(unsigned long long), + !fwrite(&events_off, sizeof(unsigned long long), 1, trace_dat_fp)) return -EIO; /* HEADER option */ - opt_id = TRACE_DAT_OPTION_HEADER; - opt_size = sizeof(unsigned long long); + opt_id = to_file_u16(pevent, TRACE_DAT_OPTION_HEADER); + opt_size = to_file_u32(pevent, sizeof(unsigned long long)); + header_off = to_file_u64(pevent, trace_dat_header_info_offset); if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) || !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) || - !fwrite(&trace_dat_header_info_offset, sizeof(unsigned long long), + !fwrite(&header_off, sizeof(unsigned long long), 1, trace_dat_fp)) return -EIO; /* FTRACE option */ - opt_id = TRACE_DAT_OPTION_FTRACE; - opt_size = sizeof(unsigned long long); + opt_id = to_file_u16(pevent, TRACE_DAT_OPTION_FTRACE); + opt_size = to_file_u32(pevent, sizeof(unsigned long long)); + ftrace_off = to_file_u64(pevent, trace_dat_ftrace_format_offset); if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) || !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) || - !fwrite(&trace_dat_ftrace_format_offset, sizeof(unsigned long long), + !fwrite(&ftrace_off, sizeof(unsigned long long), 1, trace_dat_fp)) return -EIO; /* KALLSYMS option */ - opt_id = TRACE_DAT_OPTION_KALLSYMS; - opt_size = sizeof(unsigned long long); + opt_id = to_file_u16(pevent, TRACE_DAT_OPTION_KALLSYMS); + opt_size = to_file_u32(pevent, sizeof(unsigned long long)); + kallsyms_off = to_file_u64(pevent, trace_dat_kallsyms_offset); if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) || !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) || - !fwrite(&trace_dat_kallsyms_offset, sizeof(unsigned long long), + !fwrite(&kallsyms_off, sizeof(unsigned long long), 1, trace_dat_fp)) return -EIO; /* CMDLINE option */ - opt_id = TRACE_DAT_OPTION_CMDLINE; - opt_size = sizeof(unsigned long long); + opt_id = to_file_u16(pevent, TRACE_DAT_OPTION_CMDLINE); + opt_size = to_file_u32(pevent, sizeof(unsigned long long)); + cmdline_off = to_file_u64(pevent, trace_dat_cmdline_offset); if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) || !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) || - !fwrite(&trace_dat_cmdline_offset, sizeof(unsigned long long), + !fwrite(&cmdline_off, sizeof(unsigned long long), 1, trace_dat_fp)) return -EIO; /* DONE option id - next_options_offset filled later */ - opt_id = TRACE_DAT_OPTION_DONE; - opt_size = sizeof(unsigned long long); - next_offset = 0; /* placeholder */ + opt_id = to_file_u16(pevent, TRACE_DAT_OPTION_DONE); + opt_size = to_file_u32(pevent, sizeof(unsigned long long)); + next_offset = to_file_u64(pevent, 0); trace_dat_next_options_offset = ftell(trace_dat_fp); if (!fwrite(&opt_id, sizeof(unsigned short), 1, trace_dat_fp) || @@ -560,7 +603,7 @@ int trace_dat__write_options_section1(void) /* fill section size */ end_pos = ftell(trace_dat_fp); - section_size = end_pos - payload_start; + section_size = to_file_u64(pevent, end_pos - payload_start); if (fseek(trace_dat_fp, section_size_pos, SEEK_SET) < 0 || !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp) || fseek(trace_dat_fp, end_pos, SEEK_SET) < 0) @@ -574,20 +617,21 @@ int trace_dat__write_options_section1(void) * (flyrecord section offset, clock type, page size, CPU count, * per-CPU offsets/sizes) and DONE option. */ -int trace_dat__write_options_section2(void) +int trace_dat__write_options_section2(struct tep_handle *pevent) { - unsigned short section_id = TRACE_DAT_SECTION_OPTIONS; - unsigned short flags = 0; - unsigned int string_id = STRID_OPTIONS_2; + unsigned short section_id = to_file_u16(pevent, TRACE_DAT_SECTION_OPTIONS); + unsigned short flags = to_file_u16(pevent, 0); + unsigned int string_id = to_file_u32(pevent, STRID_OPTIONS_2); unsigned long long section_size = 0; long section_size_pos; long payload_start; int cpu; - unsigned short opt_id = TRACE_DAT_OPTION_BUFFER; - unsigned int opt_size = 0; + unsigned short opt_id = to_file_u16(pevent, TRACE_DAT_OPTION_BUFFER); + unsigned int opt_size = 0; /* filled after payload written */ long opt_size_pos; unsigned long long data_offset = 0; - unsigned int page_size = (unsigned int)trace_dat_page_size; + unsigned int page_size = to_file_u32(pevent, (unsigned int)trace_dat_page_size); + unsigned int nr_cpus = to_file_u32(pevent, trace_dat_nr_cpus); const char *clock = "local"; unsigned long long next; long end_pos; @@ -600,7 +644,7 @@ int trace_dat__write_options_section2(void) return -EINVAL; /* fill done1 next offset - points to this section */ - next = ftell(trace_dat_fp); + next = to_file_u64(pevent, ftell(trace_dat_fp)); if (fseek(trace_dat_fp, trace_dat_next_options_offset + 2 + 4, SEEK_SET) < 0 || !fwrite(&next, sizeof(unsigned long long), 1, trace_dat_fp) || @@ -631,15 +675,16 @@ int trace_dat__write_options_section2(void) !fwrite("\0", 1, 1, trace_dat_fp) || !fwrite(clock, 1, strlen(clock) + 1, trace_dat_fp) || !fwrite(&page_size, sizeof(unsigned int), 1, trace_dat_fp) || - !fwrite(&trace_dat_nr_cpus, sizeof(unsigned int), 1, trace_dat_fp)) + !fwrite(&nr_cpus, sizeof(unsigned int), 1, trace_dat_fp)) return -EIO; /* per cpu: cpu_id + offset placeholder + size */ for (cpu = 0; cpu < trace_dat_nr_cpus; cpu++) { + unsigned int cpu_out = to_file_u32(pevent, cpu); cpu_offset = 0; /* filled in write_flyrecord */ - cpu_size = 0; /* filled in write_flyrecord */ + cpu_size = 0; /* filled in write_flyrecord */ - if (!fwrite(&cpu, sizeof(unsigned int), 1, trace_dat_fp)) + if (!fwrite(&cpu_out, sizeof(unsigned int), 1, trace_dat_fp)) return -EIO; buffer_opt_cpu_offsets_pos[cpu] = ftell(trace_dat_fp); if (!fwrite(&cpu_offset, sizeof(unsigned long long), 1, trace_dat_fp) || @@ -650,17 +695,17 @@ int trace_dat__write_options_section2(void) /* fill opt_size */ end_pos = ftell(trace_dat_fp); - opt_size = end_pos - opt_payload_start; - fseek(trace_dat_fp, opt_size_pos, SEEK_SET); - if (!fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp)) + opt_size = to_file_u32(pevent, end_pos - opt_payload_start); + if (fseek(trace_dat_fp, opt_size_pos, SEEK_SET) < 0 || + !fwrite(&opt_size, sizeof(unsigned int), 1, trace_dat_fp) || + fseek(trace_dat_fp, end_pos, SEEK_SET) < 0) return -EIO; - fseek(trace_dat_fp, end_pos, SEEK_SET); /* DONE id=0 */ - done_id = TRACE_DAT_OPTION_DONE; - done_size = sizeof(unsigned long long); + done_id = to_file_u16(pevent, TRACE_DAT_OPTION_DONE); + done_size = to_file_u32(pevent, sizeof(unsigned long long)); /* No additional options sections follow this one */ - next = 0; + next = to_file_u64(pevent, 0); if (!fwrite(&done_id, sizeof(unsigned short), 1, trace_dat_fp) || !fwrite(&done_size, sizeof(unsigned int), 1, trace_dat_fp) || @@ -670,24 +715,25 @@ int trace_dat__write_options_section2(void) /* fill section size */ end_pos = ftell(trace_dat_fp); - section_size = end_pos - payload_start; - fseek(trace_dat_fp, section_size_pos, SEEK_SET); - if (!fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp)) + section_size = to_file_u64(pevent, end_pos - payload_start); + if (fseek(trace_dat_fp, section_size_pos, SEEK_SET) < 0 || + !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp) || + fseek(trace_dat_fp, end_pos, SEEK_SET) < 0) return -EIO; - fseek(trace_dat_fp, end_pos, SEEK_SET); return 0; } -int trace_dat__write_flyrecord_section(void) +int trace_dat__write_flyrecord_section(struct tep_handle *pevent) { - unsigned short section_id = TRACE_DAT_SECTION_FLYRECORD; - unsigned short flags = 0; - unsigned int string_id = STRID_BUFFER_FLYRECORD; + unsigned short section_id = to_file_u16(pevent, TRACE_DAT_SECTION_FLYRECORD); + unsigned short flags = to_file_u16(pevent, 0); + unsigned int string_id = to_file_u32(pevent, STRID_BUFFER_FLYRECORD); unsigned long long section_size = 0; long section_size_pos; long flyrecord_start; + unsigned long long flyrecord_out; long after_header; long padding_needed; unsigned long long *cpu_offsets; @@ -750,7 +796,7 @@ int trace_dat__write_flyrecord_section(void) for (cpu = 0; cpu < trace_dat_nr_cpus; cpu++) { start = ftell(trace_dat_fp); - ret = trace_dat__write_cpu_dat(trace_dat_fp, cpu, &cpu_offsets[cpu]); + ret = trace_dat__write_cpu_dat(trace_dat_fp, pevent, cpu, &cpu_offsets[cpu]); if (ret < 0) { pr_err("Failed to write CPU %d data\n", cpu); @@ -762,7 +808,7 @@ int trace_dat__write_flyrecord_section(void) /* fill section size */ end_pos = ftell(trace_dat_fp); - section_size = end_pos - after_header; + section_size = to_file_u64(pevent, end_pos - after_header); if (fseek(trace_dat_fp, section_size_pos, SEEK_SET) < 0 || !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp)) { ret = -EIO; @@ -775,6 +821,8 @@ int trace_dat__write_flyrecord_section(void) /* fill cpu offsets and sizes in BUFFER option */ for (cpu = 0; cpu < trace_dat_nr_cpus; cpu++) { + cpu_offsets[cpu] = to_file_u64(pevent, cpu_offsets[cpu]); + cpu_sizes[cpu] = to_file_u64(pevent, cpu_sizes[cpu]); if (fseek(trace_dat_fp, buffer_opt_cpu_offsets_pos[cpu], SEEK_SET) < 0 || !fwrite(&cpu_offsets[cpu], sizeof(unsigned long long), 1, trace_dat_fp) || !fwrite(&cpu_sizes[cpu], sizeof(unsigned long long), 1, trace_dat_fp)) { @@ -784,8 +832,9 @@ int trace_dat__write_flyrecord_section(void) } /* fill data offset in buffer option */ + flyrecord_out = to_file_u64(pevent, flyrecord_start); if (fseek(trace_dat_fp, opt_payload_start, SEEK_SET) < 0 || - !fwrite(&flyrecord_start, sizeof(unsigned long long), 1, trace_dat_fp)) { + !fwrite(&flyrecord_out, sizeof(unsigned long long), 1, trace_dat_fp)) { ret = -EIO; goto cleanup; } @@ -795,7 +844,6 @@ int trace_dat__write_flyrecord_section(void) goto cleanup; } - cleanup: free(cpu_offsets); free(cpu_sizes); diff --git a/tools/perf/util/trace-dat.h b/tools/perf/util/trace-dat.h index 9aec37b708d4..5985083b275a 100644 --- a/tools/perf/util/trace-dat.h +++ b/tools/perf/util/trace-dat.h @@ -8,9 +8,13 @@ #define __PERF_TRACE_DAT_H #include <stdio.h> +#include <stdbool.h> +#include <event-parse.h> +#include <byteswap.h> +#include "util.h" /* trace.dat file format version */ -#define TRACE_DAT_VERSION '7' +#define TRACE_DAT_VERSION "7" /* * Section IDs for trace.dat format @@ -51,6 +55,32 @@ #define STRID_OPTIONS_2 77 #define STRID_BUFFER_FLYRECORD 85 +/* + * Set by trace.dat section writers (kallsyms, ftrace, events etc.) on + * fwrite() failure; checked by trace_convert__perf2dat() to clean up. + */ +extern bool trace_dat_write_failed; + +/* + * tep handle from the recording machine; set by trace_convert__perf2dat() + * before options/flyrecord sections are written. + */ + +static inline uint16_t to_file_u16(struct tep_handle *pevent, uint16_t val) +{ + return tep_read_number(pevent, &val, 2); +} + +static inline uint32_t to_file_u32(struct tep_handle *pevent, uint32_t val) +{ + return tep_read_number(pevent, &val, 4); +} + +static inline uint64_t to_file_u64(struct tep_handle *pevent, uint64_t val) +{ + return tep_read_number(pevent, &val, 8); +} + struct perf_session; extern FILE *trace_dat_fp; @@ -75,9 +105,9 @@ int trace_dat__collect_cpu_event(int cpu, unsigned long long ts, void trace_dat__free_cpu_buffers(void); /* write trace.dat file sections */ -int trace_dat__write_options_section1(void); -int trace_dat__write_options_section2(void); -int trace_dat__write_flyrecord_section(void); -int trace_dat__write_strings_section(void); +int trace_dat__write_options_section1(struct tep_handle *pevent); +int trace_dat__write_options_section2(struct tep_handle *pevent); +int trace_dat__write_flyrecord_section(struct tep_handle *pevent); +int trace_dat__write_strings_section(struct tep_handle *pevent); #endif /* __PERF_TRACE_DAT_H */ diff --git a/tools/perf/util/trace-event-read.c b/tools/perf/util/trace-event-read.c index ecbbb93f0185..7e38e9c8f507 100644 --- a/tools/perf/util/trace-event-read.c +++ b/tools/perf/util/trace-event-read.c @@ -19,6 +19,7 @@ #include "trace-event.h" #include "debug.h" #include "util.h" +#include "trace-dat.h" static int input_fd; @@ -145,10 +146,9 @@ static char *read_string(void) static int read_proc_kallsyms(struct tep_handle *pevent) { unsigned int size; + char *buf; size = read4(pevent); - if (!size) - return 0; /* * Just skip it, now that we configure libtraceevent to use the * tools/perf/ symbol resolver. @@ -160,11 +160,59 @@ static int read_proc_kallsyms(struct tep_handle *pevent) * payload", so that older tools can continue reading it and interpret * it as "no kallsyms payload is present". */ - lseek(input_fd, size, SEEK_CUR); + /* Write kallsyms section with empty payload if no data */ + if (!size) { + if (trace_dat_fp && !trace_dat_write_failed) { + unsigned short section_id = to_file_u16(pevent, TRACE_DAT_SECTION_KALLSYMS); + unsigned short flags = to_file_u16(pevent, 0); + unsigned int string_id = to_file_u32(pevent, STRID_KALLSYMS); + unsigned long long section_size = to_file_u64(pevent, sizeof(unsigned int)); + unsigned int kallsyms_data = to_file_u32(pevent, 0); + + trace_dat_kallsyms_offset = ftell(trace_dat_fp); + if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp) || + !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp) || + !fwrite(&kallsyms_data, sizeof(unsigned int), 1, trace_dat_fp)) { + pr_warning("Failed to write trace.dat kallsyms section\n"); + trace_dat_write_failed = true; + } + } + return 0; + } + buf = malloc(size); + if (buf == NULL) + return -1; + if (do_read(buf, size) < 0) { + free(buf); + return -1; + } trace_data_size += size; + /* Write kallsyms section with data */ + if (trace_dat_fp && !trace_dat_write_failed) { + unsigned short section_id = to_file_u16(pevent, TRACE_DAT_SECTION_KALLSYMS); + unsigned int string_id = to_file_u32(pevent, STRID_KALLSYMS); + unsigned long long section_size = to_file_u64(pevent, sizeof(unsigned int) + size); + unsigned short flags = to_file_u16(pevent, 0); + unsigned int size_out = to_file_u32(pevent, size); + + trace_dat_kallsyms_offset = ftell(trace_dat_fp); + if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp) || + !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp) || + !fwrite(&size_out, sizeof(unsigned int), 1, trace_dat_fp) || + !fwrite(buf, 1, size, trace_dat_fp)) { + pr_warning("Failed to write trace.dat kallsyms section\n"); + trace_dat_write_failed = true; + } + } + free(buf); return 0; } + static int read_ftrace_printk(struct tep_handle *pevent) { unsigned int size; @@ -195,6 +243,13 @@ static int read_ftrace_printk(struct tep_handle *pevent) static int read_header_files(struct tep_handle *pevent) { unsigned long long size; + unsigned long long header_page_size; + unsigned long long header_event_size; + char *header_event; + unsigned short section_id; + unsigned short flags; + unsigned int string_id; + unsigned long long section_size; char *header_page; char buf[BUFSIZ]; int ret = 0; @@ -209,6 +264,7 @@ static int read_header_files(struct tep_handle *pevent) size = read8(pevent); + header_page_size = size; header_page = malloc(size); if (header_page == NULL) return -1; @@ -227,19 +283,63 @@ static int read_header_files(struct tep_handle *pevent) */ tep_set_long_size(pevent, tep_get_header_page_size(pevent)); } - free(header_page); - if (do_read(buf, 13) < 0) + if (do_read(buf, 13) < 0) { + free(header_page); return -1; + } if (memcmp(buf, "header_event", 13) != 0) { pr_debug("did not read header event"); + free(header_page); return -1; } size = read8(pevent); - skip(size); + if (trace_dat_fp && !trace_dat_write_failed) { + unsigned long long header_page_size_out; + unsigned long long header_event_size_out; + + header_event_size = size; + header_event = malloc(size); + if (header_event == NULL) { + free(header_page); + return -1; + } + if (do_read(header_event, size) < 0) { + free(header_page); + free(header_event); + return -1; + } + /* Write header_page and header_event to trace.dat */ + section_id = to_file_u16(pevent, TRACE_DAT_SECTION_HEADER); + flags = to_file_u16(pevent, 0); + string_id = to_file_u32(pevent, STRID_HEADERS); + section_size = to_file_u64(pevent, 12 + 8 + header_page_size + + 13 + 8 + header_event_size); + header_page_size_out = to_file_u64(pevent, header_page_size); + header_event_size_out = to_file_u64(pevent, header_event_size); + + trace_dat_header_info_offset = ftell(trace_dat_fp); + if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp) || + !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp) || + !fwrite("header_page\0", 1, 12, trace_dat_fp) || + !fwrite(&header_page_size_out, sizeof(unsigned long long), 1, trace_dat_fp) || + !fwrite(header_page, 1, header_page_size, trace_dat_fp) || + !fwrite("header_event\0", 1, 13, trace_dat_fp) || + !fwrite(&header_event_size_out, sizeof(unsigned long long), 1, trace_dat_fp) || + !fwrite(header_event, 1, header_event_size, trace_dat_fp)) { + pr_warning("Failed to write trace.dat header section\n"); + trace_dat_write_failed = true; + } + free(header_event); + } else { + skip(size); + } + free(header_page); return ret; } @@ -259,6 +359,15 @@ static int read_ftrace_file(struct tep_handle *pevent, unsigned long long size) pr_debug("error reading ftrace file.\n"); goto out; } + if (trace_dat_fp && !trace_dat_write_failed) { + unsigned long long size_out = to_file_u64(pevent, size); + + if (!fwrite(&size_out, sizeof(unsigned long long), 1, trace_dat_fp) || + !fwrite(buf, 1, size, trace_dat_fp)) { + pr_warning("Failed to write trace.dat ftrace event formats section\n"); + trace_dat_write_failed = true; + } + } ret = parse_ftrace_file(pevent, buf, size); if (ret < 0) @@ -283,6 +392,15 @@ static int read_event_file(struct tep_handle *pevent, char *sys, ret = do_read(buf, size); if (ret < 0) goto out; + if (trace_dat_fp && !trace_dat_write_failed) { + unsigned long long size_out = to_file_u64(pevent, size); + + if (!fwrite(&size_out, sizeof(unsigned long long), 1, trace_dat_fp) || + !fwrite(buf, 1, size, trace_dat_fp)) { + pr_warning("Failed to write trace.dat event formats section\n"); + trace_dat_write_failed = true; + } + } ret = parse_event_file(pevent, buf, size, sys); if (ret < 0) @@ -298,8 +416,39 @@ static int read_ftrace_files(struct tep_handle *pevent) int count; int i; int ret; + long section_size_pos = 0; + long count_pos = 0; + unsigned long long section_size = 0; + long end_pos; count = read4(pevent); + /* Write ftrace formats section to trace.dat output file */ + if (trace_dat_fp && !trace_dat_write_failed) { + unsigned short section_id = to_file_u16(pevent, TRACE_DAT_SECTION_FTRACE); + unsigned short flags = to_file_u16(pevent, 0); + unsigned int string_id = to_file_u32(pevent, STRID_FTRACE_FORMATS); + unsigned int count_out = to_file_u32(pevent, count); + + section_size = to_file_u64(pevent, 0); + trace_dat_ftrace_format_offset = ftell(trace_dat_fp); + + if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp)) { + pr_warning("Failed to write trace.dat ftrace event formats section\n"); + trace_dat_write_failed = true; + } + section_size_pos = ftell(trace_dat_fp); + if (!fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp)) { + pr_warning("Failed to write trace.dat ftrace event formats section\n"); + trace_dat_write_failed = true; + } + count_pos = ftell(trace_dat_fp); + if (!fwrite(&count_out, sizeof(unsigned int), 1, trace_dat_fp)) { + pr_warning("Failed to write trace.dat ftrace event formats section\n"); + trace_dat_write_failed = true; + } + } for (i = 0; i < count; i++) { size = read8(pevent); @@ -307,6 +456,18 @@ static int read_ftrace_files(struct tep_handle *pevent) if (ret) return ret; } + /* Fill in section size after writing all ftrace files */ + if (trace_dat_fp && !trace_dat_write_failed) { + end_pos = ftell(trace_dat_fp); + section_size = to_file_u64(pevent, end_pos - count_pos); + if (fseek(trace_dat_fp, section_size_pos, SEEK_SET) < 0 || + !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp) || + fseek(trace_dat_fp, end_pos, SEEK_SET) < 0) { + pr_warning("Failed to write trace.dat ftrace event formats section\n"); + trace_dat_write_failed = true; + } + } + return 0; } @@ -318,8 +479,38 @@ static int read_event_files(struct tep_handle *pevent) int count; int i,x; int ret; + long section_size_pos = 0; + long sys_count_pos = 0; + unsigned long long section_size = 0; + long end_pos; systems = read4(pevent); + /* Write event formats section to trace.dat output file */ + if (trace_dat_fp && !trace_dat_write_failed) { + unsigned short section_id = to_file_u16(pevent, TRACE_DAT_SECTION_EVENTS); + unsigned short flags = to_file_u16(pevent, 0); + unsigned int string_id = to_file_u32(pevent, STRID_EVENT_FORMATS); + unsigned int systems_out = to_file_u32(pevent, systems); + + section_size = to_file_u64(pevent, 0); + trace_dat_events_format_offset = ftell(trace_dat_fp); + if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp)) { + pr_warning("Failed to write trace.dat event formats section\n"); + trace_dat_write_failed = true; + } + section_size_pos = ftell(trace_dat_fp); + if (!fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp)) { + pr_warning("Failed to write trace.dat event formats section\n"); + trace_dat_write_failed = true; + } + sys_count_pos = ftell(trace_dat_fp); + if (!fwrite(&systems_out, sizeof(unsigned int), 1, trace_dat_fp)) { + pr_warning("Failed to write trace.dat event formats section\n"); + trace_dat_write_failed = true; + } + } for (i = 0; i < systems; i++) { sys = read_string(); @@ -327,6 +518,15 @@ static int read_event_files(struct tep_handle *pevent) return -1; count = read4(pevent); + if (trace_dat_fp && !trace_dat_write_failed) { + unsigned int count_out = to_file_u32(pevent, count); + + if (!fwrite(sys, 1, strlen(sys) + 1, trace_dat_fp) || + !fwrite(&count_out, sizeof(unsigned int), 1, trace_dat_fp)) { + pr_warning("Failed to write trace.dat event formats section\n"); + trace_dat_write_failed = true; + } + } for (x=0; x < count; x++) { size = read8(pevent); @@ -338,6 +538,18 @@ static int read_event_files(struct tep_handle *pevent) } free(sys); } + /* Fill in section size after writing all event files */ + if (trace_dat_fp && !trace_dat_write_failed) { + end_pos = ftell(trace_dat_fp); + section_size = to_file_u64(pevent, end_pos - sys_count_pos); + fseek(trace_dat_fp, section_size_pos, SEEK_SET); + if (!fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp)) { + pr_warning("Failed to write trace.dat event formats section\n"); + trace_dat_write_failed = true; + } + fseek(trace_dat_fp, end_pos, SEEK_SET); + } + return 0; } @@ -349,8 +561,28 @@ static int read_saved_cmdline(struct tep_handle *pevent) /* it can have 0 size */ size = read8(pevent); - if (!size) + /* Write cmdlines section with empty payload if no data */ + if (!size) { + if (trace_dat_fp && !trace_dat_write_failed) { + unsigned short section_id = to_file_u16(pevent, TRACE_DAT_SECTION_CMDLINE); + unsigned short flags = to_file_u16(pevent, 0); + unsigned int string_id = to_file_u32(pevent, STRID_CMDLINES); + unsigned long long section_size = + to_file_u64(pevent, sizeof(unsigned long long)); + unsigned long long section_data = to_file_u64(pevent, 0); + + trace_dat_cmdline_offset = ftell(trace_dat_fp); + if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp) || + !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp) || + !fwrite(§ion_data, sizeof(unsigned long long), 1, trace_dat_fp)) { + pr_warning("Failed to write trace.dat cmdlines section\n"); + trace_dat_write_failed = true; + } + } return 0; + } buf = malloc(size + 1); if (buf == NULL) { @@ -363,6 +595,27 @@ static int read_saved_cmdline(struct tep_handle *pevent) pr_debug("error reading saved cmdlines\n"); goto out; } + /* Write cmdlines section with data */ + if (trace_dat_fp && !trace_dat_write_failed) { + unsigned short section_id = to_file_u16(pevent, TRACE_DAT_SECTION_CMDLINE); + unsigned short flags = to_file_u16(pevent, 0); + unsigned int string_id = to_file_u32(pevent, STRID_CMDLINES); + unsigned long long section_size = + to_file_u64(pevent, sizeof(unsigned long long) + size); + unsigned long long size_out = to_file_u64(pevent, size); + + trace_dat_cmdline_offset = ftell(trace_dat_fp); + if (!fwrite(§ion_id, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&flags, sizeof(unsigned short), 1, trace_dat_fp) || + !fwrite(&string_id, sizeof(unsigned int), 1, trace_dat_fp) || + !fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp) || + !fwrite(&size_out, sizeof(unsigned long long), 1, trace_dat_fp) || + !fwrite(buf, 1, size, trace_dat_fp)) { + pr_warning("Failed to write trace.dat cmdlines section\n"); + trace_dat_write_failed = true; + } + } + buf[ret] = '\0'; parse_saved_cmdline(pevent, buf, size); @@ -387,6 +640,7 @@ ssize_t trace_report(int fd, struct trace_event *tevent, bool __repipe) int file_page_size; struct tep_handle *pevent = NULL; int err; + char magic_buf[10]; repipe = __repipe; input_fd = fd; @@ -398,12 +652,17 @@ ssize_t trace_report(int fd, struct trace_event *tevent, bool __repipe) return -1; } + if (trace_dat_fp) + memcpy(magic_buf, buf, 3); + if (do_read(buf, 7) < 0) return -1; if (memcmp(buf, "tracing", 7) != 0) { pr_debug("not a trace file (missing 'tracing' tag)"); return -1; } + if (trace_dat_fp) + memcpy(magic_buf + 3, buf, 7); version = read_string(); if (version == NULL) @@ -440,6 +699,34 @@ ssize_t trace_report(int fd, struct trace_event *tevent, bool __repipe) tep_set_long_size(pevent, file_long_size); tep_set_page_size(pevent, file_page_size); + /* Write initial file header to trace.dat */ + if (trace_dat_fp && !trace_dat_write_failed) { + unsigned char endian = file_bigendian; + unsigned char long_size = file_long_size; + unsigned int page_size = to_file_u32(pevent, file_page_size); + unsigned long long placeholder = to_file_u64(pevent, 0); + + if (!fwrite(magic_buf, 1, 10, trace_dat_fp) || /* magic + "tracing" */ + !fwrite(TRACE_DAT_VERSION, 1, 2, trace_dat_fp) || + !fwrite(&endian, 1, 1, trace_dat_fp) || + !fwrite(&long_size, 1, 1, trace_dat_fp) || + !fwrite(&page_size, sizeof(unsigned int), 1, trace_dat_fp) || + !fwrite("none", 1, 4, trace_dat_fp) || + !fwrite("\0", 1, 1, trace_dat_fp) || + !fwrite("\0", 1, 1, trace_dat_fp)) { + pr_warning("Failed to write trace.dat initalfile header\n"); + trace_dat_write_failed = true; + } + + if (!trace_dat_write_failed) { + trace_dat_options_offset = ftell(trace_dat_fp); + if (!fwrite(&placeholder, sizeof(unsigned long long), 1, trace_dat_fp)) { + pr_warning("Failed to write trace.dat initial file header\n"); + trace_dat_write_failed = true; + } + } + } + err = read_header_files(pevent); if (err) goto out; @@ -460,6 +747,12 @@ ssize_t trace_report(int fd, struct trace_event *tevent, bool __repipe) if (err) goto out; } + /* Write strings section to trace.dat output file */ + if (trace_dat_fp && !trace_dat_write_failed) { + err = trace_dat__write_strings_section(pevent); + if (err) + goto out; + } size = trace_data_size; repipe = false; -- 2.47.1 ^ permalink raw reply related [flat|nested] 14+ messages in thread
* Re: [RFC PATCH v3 2/5] perf/trace-event: Write trace.dat metadata sections during parsing 2026-08-03 14:59 ` [RFC PATCH v3 2/5] perf/trace-event: Write trace.dat metadata sections during parsing Tanushree Shah @ 2026-08-03 15:14 ` sashiko-bot 0 siblings, 0 replies; 14+ messages in thread From: sashiko-bot @ 2026-08-03 15:14 UTC (permalink / raw) To: Tanushree Shah; +Cc: linux-perf-users Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider: New issues: - [Medium] Missing checks for `fseek` errors when patching section sizes can silently corrupt the output trace file. - [Low] Missing explicit inclusion of `<stdint.h>` in `trace-dat.h`. Pre-existing issues: - [Critical] Integer overflow in `malloc(size + 1)` leads to a massive heap buffer overflow. -- --- Patch [2]: [RFC PATCH v3 2/5] perf/trace-event: Write trace.dat metadata sections during parsing --- commit 9121c22ee8da1e40687be01e1d9fb0edd88a10cd Author: Tanushree Shah <tshah@linux.ibm.com> perf/trace-event: Write trace.dat metadata sections during parsing This commit modifies the perf trace event parsing logic to write trace.dat compatible metadata sections when trace_dat_fp is set. It introduces byte-order helpers and adds support for writing FTRACE EVENT FORMATS, KALLSYMS, CMDLINES, and STRINGS sections. > diff --git a/tools/perf/util/trace-dat.h b/tools/perf/util/trace-dat.h > index 9aec37b708d41..5985083b275ac 100644 > --- a/tools/perf/util/trace-dat.h > +++ b/tools/perf/util/trace-dat.h > @@ -8,9 +8,13 @@ > #define __PERF_TRACE_DAT_H > > #include <stdio.h> > +#include <stdbool.h> > +#include <event-parse.h> > +#include <byteswap.h> > +#include "util.h" [Severity: Low] This isn't a bug, but could the lack of an explicit <stdint.h> inclusion here cause build failures on musl libc? The file uses uint16_t, uint32_t, and uint64_t, and according to subsystem guidelines, all files using POSIX libc types must explicitly include the declaring header. [ ... ] > --- a/tools/perf/util/trace-event-read.c > +++ b/tools/perf/util/trace-event-read.c [ ... ] > @@ -338,6 +538,18 @@ static int read_event_files(struct tep_handle *pevent) > } > free(sys); > } > + /* Fill in section size after writing all event files */ > + if (trace_dat_fp && !trace_dat_write_failed) { > + end_pos = ftell(trace_dat_fp); > + section_size = to_file_u64(pevent, end_pos - sys_count_pos); > + fseek(trace_dat_fp, section_size_pos, SEEK_SET); > + if (!fwrite(§ion_size, sizeof(unsigned long long), 1, trace_dat_fp)) { > + pr_warning("Failed to write trace.dat event formats section\n"); > + trace_dat_write_failed = true; > + } > + fseek(trace_dat_fp, end_pos, SEEK_SET); [Severity: Medium] Can a failed fseek() here silently corrupt the output trace.dat file? If the first fseek() fails due to an underlying I/O error, the subsequent fwrite() would overwrite data at the wrong location. Since the second fseek() is also unchecked, subsequent writes would also land at invalid offsets. > + } > + > return 0; > } [ ... ] > @@ -349,8 +561,28 @@ static int read_saved_cmdline(struct tep_handle *pevent) > > /* it can have 0 size */ > size = read8(pevent); > - if (!size) > + /* Write cmdlines section with empty payload if no data */ > + if (!size) { > + if (trace_dat_fp && !trace_dat_write_failed) { [ ... ] > + } > return 0; > + } > > buf = malloc(size + 1); [Severity: Critical] This is a pre-existing issue, but could this malloc(size + 1) trigger a heap buffer overflow regression when parsing an untrusted perf.data file? For example, in read_ftrace_printk(), size is a 32-bit unsigned int. If size is UINT_MAX, size + 1 deterministically wraps to 0. This results in a 0-byte allocation, into which do_read() subsequently attempts to read UINT_MAX bytes. A similar issue exists here in read_saved_cmdline() on 32-bit architectures where size_t is 32-bit. > if (buf == NULL) { -- Sashiko AI review · https://sashiko.dev/#/patchset/20260803145958.299956-1-tshah@linux.ibm.com?part=2 ^ permalink raw reply [flat|nested] 14+ messages in thread
* [RFC PATCH v3 3/5] perf data-convert: Add perf.data to trace.dat conversion backend 2026-08-03 14:59 [RFC PATCH v3 0/5] Add perf.data tracepoint events to trace.dat conversion Tanushree Shah 2026-08-03 14:59 ` [RFC PATCH v3 1/5] perf/trace-dat: Add trace.dat export infrastructure Tanushree Shah 2026-08-03 14:59 ` [RFC PATCH v3 2/5] perf/trace-event: Write trace.dat metadata sections during parsing Tanushree Shah @ 2026-08-03 14:59 ` Tanushree Shah 2026-08-03 15:11 ` sashiko-bot 2026-08-03 14:59 ` [RFC PATCH v3 4/5] perf data: Add --to-trace-dat option for converting perf.data tracepoint events into trace.dat format Tanushree Shah ` (2 subsequent siblings) 5 siblings, 1 reply; 14+ messages in thread From: Tanushree Shah @ 2026-08-03 14:59 UTC (permalink / raw) To: acme, jolsa, adrian.hunter, vmolnaro, mpetlan, tmricht, maddy, irogers, namhyung Cc: linux-perf-users, linuxppc-dev, atrajeev, hbathini, Tejas.Manhas1, Tanushree.Shah, rostedt, Tanushree Shah Add data-convert-trace.c implementing trace_convert__perf2dat() to convert perf.data tracepoint events to trace.dat format. process_sample_event() is invoked for each PERF_TYPE_TRACEPOINT sample during perf_session__process_events(), storing raw event bytes per-cpu via trace_dat__collect_cpu_event(). Once all samples are collected: - trace_dat__write_options_section1() writes the OPTIONS section with CPUCOUNT, TRACECLOCK, HEADER_INFO, FTRACE_EVENTS, EVENT_FORMATS, KALLSYMS, CMDLINES and DONE options. - trace_dat__write_options_section2() writes the OPTIONS section with BUFFER option holding per-cpu data offset placeholders and the DONE option. - trace_dat__write_flyrecord_section() builds ring buffer pages per-cpu and patches BUFFER option with final offsets and sizes, Per-cpu buffers are sized to tep_get_page_size() from the session tep handle and released on all exit paths. Register .attr, .feature, and .tracing_data callbacks on the tool, delegating to the standard perf_event__process_attr(), perf_event__process_feature(), and perf_event__process_tracing_data() helpers, so these records are processed as they arrive in the stream instead of being left as default no-op stubs. Defer per-CPU buffer initialisation to the first call of the .sample callback instead of doing it eagerly after session open. By this point, pipe-mode ordering guarantees (record__synthesize()) that attr, feature, and tracing_data records have already been written and processed, so nr_cpus_online and the tep page size are guaranteed to be populated from the recording machine's own stream data. This is required because the converting host may differ in architecture, CPU count, or page size from the machine that recorded the data, so these values must never be taken from local sysconf() or similar. The output path honours --force: without it, the destination file is opened with O_CREAT|O_EXCL so an existing file is detected atomically and conversion fails with -EEXIST rather than silently overwriting prior output. Signed-off-by: Tanushree Shah <tshah@linux.ibm.com> --- tools/perf/util/Build | 1 + tools/perf/util/data-convert-trace.c | 241 +++++++++++++++++++++++++++ tools/perf/util/data-convert.h | 4 + 3 files changed, 246 insertions(+) create mode 100644 tools/perf/util/data-convert-trace.c diff --git a/tools/perf/util/Build b/tools/perf/util/Build index 246d05ce25cd..b95191226f21 100644 --- a/tools/perf/util/Build +++ b/tools/perf/util/Build @@ -235,6 +235,7 @@ ifeq ($(CONFIG_LIBTRACEEVENT),y) endif perf-util-y += data-convert-json.o +perf-util-$(CONFIG_LIBTRACEEVENT) += data-convert-trace.o perf-util-y += scripting-engines/ diff --git a/tools/perf/util/data-convert-trace.c b/tools/perf/util/data-convert-trace.c new file mode 100644 index 000000000000..8dbdb2c9caa4 --- /dev/null +++ b/tools/perf/util/data-convert-trace.c @@ -0,0 +1,241 @@ +// SPDX-License-Identifier: GPL-2.0-or-later +/* + * Copyright 2026, IBM Corporation + * Author: Tanushree Shah <tshah@linux.ibm.com> + * + * data-convert-trace.c + * + * Implements perf.data to trace.dat format conversion for tracepoint events. + */ + +#include <errno.h> +#include <inttypes.h> +#include <fcntl.h> +#include <string.h> +#include <unistd.h> +#include <linux/compiler.h> +#include <linux/err.h> + +#include "data-convert.h" +#include "session.h" +#include "evsel.h" +#include "tool.h" +#include "debug.h" +#include "trace-dat.h" +#include "trace-event.h" +#include "event.h" +#include "sample.h" +#include "evlist.h" + +struct trace_convert { + struct perf_tool tool; + u64 events_count; +}; + +/* Session handle and init flag used for lazy CPU buffer init in pipe mode */ +static struct perf_session *trace_dat_session; +static bool cpu_buffers_initialized; + +/* Store raw tracepoint event data in per-cpu buffer for trace.dat flyrecord */ +static int process_sample_event(const struct perf_tool *tool, + union perf_event *event __maybe_unused, + struct perf_sample *sample, + struct machine *machine __maybe_unused) +{ + struct trace_convert *tc = container_of(tool, struct trace_convert, tool); + struct evsel *evsel = sample->evsel; + + /* Only process tracepoint events */ + if (!trace_dat_fp || sample->raw_size == 0 || + !evsel || + evsel->core.attr.type != PERF_TYPE_TRACEPOINT) + return 0; + + /* + * In pipe mode, CPU count and page size arrive via feature/tracing_data + * records before the first sample; initialize buffers lazily on first sample. + */ + if (!cpu_buffers_initialized) { + int nr_cpus = trace_dat_session->header.env.nr_cpus_online; + + if (trace_dat_session->tevent.pevent) + trace_dat_page_size = tep_get_page_size(trace_dat_session->tevent.pevent); + + /* + * nr_cpus and page size MUST come from the recording + * machine's own stream data (feature/tracing_data records + * never from local sysconf() - the converting host may be + * a different arch/cpu-count/page-size than the machine + * that recorded the data. Pipe-mode ordering guarantees + * (record__synthesize()) that attr -> feature -> + * tracing_data are always written before the first sample, + * so by the time we're in this callback, both values + * should already be populated by process_feature()/ + * process_tracing_data(). + */ + if (nr_cpus <= 0 || trace_dat_page_size == 0) { + pr_err("Malformed pipe stream: missing feature/tracing_data before first sample\n"); + return -EINVAL; + } + + if (trace_dat__init_cpu_buffers(nr_cpus) < 0) { + pr_err("Failed to initialize CPU buffers\n"); + return -ENOMEM; + } + cpu_buffers_initialized = true; + } + + if (trace_dat__collect_cpu_event(sample->cpu, sample->time, + sample->raw_data, sample->raw_size) < 0) { + pr_err("Failed to collect CPU event\n"); + return -ENOMEM; + } + tc->events_count++; + + return 0; +} + +/* Process event attributes for pipe mode */ +static int process_attr(const struct perf_tool *tool __maybe_unused, + union perf_event *event, + struct evlist **pevlist) +{ + return perf_event__process_attr(tool, event, pevlist); +} + +/* Process feature records for pipe mode */ +static int process_feature(const struct perf_tool *tool __maybe_unused, + struct perf_session *session, + union perf_event *event) +{ + return perf_event__process_feature(tool, session, event); +} + +/* Process tracing data for pipe mode */ +static int process_tracing_data(const struct perf_tool *tool __maybe_unused, + struct perf_session *session, + union perf_event *event) +{ + return perf_event__process_tracing_data(tool, session, event); +} + +/* Convert perf.data tracepoint events to trace.dat format */ +int trace_convert__perf2dat(const char *input, const char *to_trace, + struct perf_data_convert_opts *opts) +{ + struct perf_session *session; + struct trace_convert tc = { + .events_count = 0, + }; + struct perf_data data = { + .path = input, + .mode = PERF_DATA_MODE_READ, + .force = opts->force, + }; + int ret = -EINVAL; + + cpu_buffers_initialized = false; + + /* Initialize tool with all required callbacks */ + perf_tool__init(&tc.tool, /*ordered_events=*/true); + tc.tool.sample = process_sample_event; + tc.tool.attr = process_attr; + tc.tool.feature = process_feature; + tc.tool.tracing_data = process_tracing_data; + + if (!opts->force) { + int fd = open(to_trace, O_WRONLY | O_CREAT | O_EXCL, 0644); + + if (fd < 0) { + if (errno == EEXIST) + pr_err("Output file '%s' already exists. Use --force to overwrite.\n", + to_trace); + else + pr_err("Failed to open output file '%s': %s\n", + to_trace, strerror(errno)); + return -errno; + } + trace_dat_fp = fdopen(fd, "wb"); + if (!trace_dat_fp) { + int err = errno; + + close(fd); + pr_err("Failed to open output file '%s': %s\n", + to_trace, strerror(err)); + return -err; + } + } else { + trace_dat_fp = fopen(to_trace, "wb"); + if (!trace_dat_fp) { + pr_err("Failed to open output file: %s\n", to_trace); + return -EINVAL; + } + } + + /* Open perf.data session - this writes trace.dat metadata sections */ + session = perf_session__new(&data, &tc.tool); + if (IS_ERR(session)) { + pr_err("Failed to open perf.data file\n"); + ret = PTR_ERR(session); + goto out_close; + } + + /* Stash session for lazy CPU buffer init on first sample (pipe and normal mode) */ + trace_dat_session = session; + + /* Process all events - collects raw data per-cpu */ + ret = perf_session__process_events(session); + if (ret < 0) { + pr_err("Failed to process events\n"); + goto out_delete; + } + + /* Skip file creation if no tracepoint events found */ + if (tc.events_count == 0) { + pr_warning("No tracepoint events found in '%s', skipping trace.dat creation\n", + input); + ret = -EINVAL; + goto out_delete; + } + + /* Write trace.dat options and flyrecord sections */ + if (trace_dat__write_options_section1(session->tevent.pevent) < 0 || + trace_dat_write_failed) { + pr_err("Failed to write options section1\n"); + ret = -EIO; + goto out_delete; + } + if (trace_dat__write_options_section2(session->tevent.pevent) < 0 || + trace_dat_write_failed) { + pr_err("Failed to write options section2\n"); + ret = -EIO; + goto out_delete; + } + if (trace_dat__write_flyrecord_section(session->tevent.pevent) < 0 || + trace_dat_write_failed) { + pr_err("Failed to write flyrecord section\n"); + ret = -EIO; + goto out_delete; + } + + pr_info("[ perf data convert: Converted '%s' into trace.dat format '%s' ]\n", + input, to_trace); + pr_info("[ perf data convert: Converted %llu events ]\n", + (unsigned long long)tc.events_count); + + ret = 0; + +out_delete: + if (cpu_buffers_initialized) + trace_dat__free_cpu_buffers(); + perf_session__delete(session); + trace_dat_session = NULL; +out_close: + if (trace_dat_fp) { + fclose(trace_dat_fp); + trace_dat_fp = NULL; + } + if (ret != 0) + unlink(to_trace); + return ret; +} diff --git a/tools/perf/util/data-convert.h b/tools/perf/util/data-convert.h index a96240f15671..f041c2325226 100644 --- a/tools/perf/util/data-convert.h +++ b/tools/perf/util/data-convert.h @@ -19,4 +19,8 @@ int bt_convert__perf2ctf(const char *input_name, const char *to_ctf, int bt_convert__perf2json(const char *input_name, const char *to_ctf, struct perf_data_convert_opts *opts); +#ifdef HAVE_LIBTRACEEVENT +int trace_convert__perf2dat(const char *input, const char *to_trace, + struct perf_data_convert_opts *opts); +#endif /* HAVE_LIBTRACEEVENT */ #endif /* __DATA_CONVERT_H */ -- 2.47.1 ^ permalink raw reply related [flat|nested] 14+ messages in thread
* Re: [RFC PATCH v3 3/5] perf data-convert: Add perf.data to trace.dat conversion backend 2026-08-03 14:59 ` [RFC PATCH v3 3/5] perf data-convert: Add perf.data to trace.dat conversion backend Tanushree Shah @ 2026-08-03 15:11 ` sashiko-bot 0 siblings, 0 replies; 14+ messages in thread From: sashiko-bot @ 2026-08-03 15:11 UTC (permalink / raw) To: Tanushree Shah; +Cc: linux-perf-users Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider: - [Medium] Missing inline fallback stub for `trace_convert__perf2dat` when `HAVE_LIBTRACEEVENT` is undefined. - [Medium] Missing POSIX header `<stdio.h>` for file I/O operations. - [Medium] Resource leak: orphaned file left on the filesystem if `fdopen()` fails. -- --- Patch [3]: [RFC PATCH v3 3/5] perf data-convert: Add perf.data to trace.dat conversion backend --- commit 25ab95b73fc4726a00bca4e0429ad24d8f04129d Author: Tanushree Shah <tshah@linux.ibm.com> perf data-convert: Add perf.data to trace.dat conversion backend Add data-convert-trace.c implementing trace_convert__perf2dat() to convert perf.data tracepoint events to trace.dat format. > diff --git a/tools/perf/util/data-convert-trace.c b/tools/perf/util/data-convert-trace.c > new file mode 100644 > index 0000000000000..8dbdb2c9caa4f > --- /dev/null > +++ b/tools/perf/util/data-convert-trace.c > @@ -0,0 +1,241 @@ > +// SPDX-License-Identifier: GPL-2.0-or-later [ ... ] > +#include <errno.h> > +#include <inttypes.h> > +#include <fcntl.h> > +#include <string.h> > +#include <unistd.h> [Severity: Medium] Does this file also need to include <stdio.h>? The code heavily relies on FILE pointers and stdio routines like fopen(), fdopen(), and fclose(). Without the explicit include, builds might fail on musl libc due to relying on implicit inclusions. > +#include <linux/compiler.h> > +#include <linux/err.h> [ ... ] > +int trace_convert__perf2dat(const char *input, const char *to_trace, > + struct perf_data_convert_opts *opts) > +{ [ ... ] > + if (!opts->force) { > + int fd = open(to_trace, O_WRONLY | O_CREAT | O_EXCL, 0644); > + > + if (fd < 0) { [ ... ] > + return -errno; > + } > + trace_dat_fp = fdopen(fd, "wb"); > + if (!trace_dat_fp) { > + int err = errno; > + > + close(fd); > + pr_err("Failed to open output file '%s': %s\n", > + to_trace, strerror(err)); > + return -err; [Severity: Medium] Can this error path leave an orphaned file on the filesystem? By returning directly instead of jumping to the out_close label where unlink(to_trace) is called, the newly created 0-byte file isn't cleaned up. This might break subsequent runs without --force since O_EXCL would fail with EEXIST. > + } [ ... ] > diff --git a/tools/perf/util/data-convert.h b/tools/perf/util/data-convert.h > index a96240f156712..f041c23252263 100644 > --- a/tools/perf/util/data-convert.h > +++ b/tools/perf/util/data-convert.h > @@ -19,4 +19,8 @@ int bt_convert__perf2ctf(const char *input_name, const char *to_ctf, > int bt_convert__perf2json(const char *input_name, const char *to_ctf, > struct perf_data_convert_opts *opts); > > +#ifdef HAVE_LIBTRACEEVENT > +int trace_convert__perf2dat(const char *input, const char *to_trace, > + struct perf_data_convert_opts *opts); > +#endif /* HAVE_LIBTRACEEVENT */ [Severity: Medium] Should this declaration include a corresponding #else block with a dummy inline fallback? When libtraceevent is disabled, omitting the stub could cause compilation failures if callers do not wrap their usage in #ifdef blocks. -- Sashiko AI review · https://sashiko.dev/#/patchset/20260803145958.299956-1-tshah@linux.ibm.com?part=3 ^ permalink raw reply [flat|nested] 14+ messages in thread
* [RFC PATCH v3 4/5] perf data: Add --to-trace-dat option for converting perf.data tracepoint events into trace.dat format 2026-08-03 14:59 [RFC PATCH v3 0/5] Add perf.data tracepoint events to trace.dat conversion Tanushree Shah ` (2 preceding siblings ...) 2026-08-03 14:59 ` [RFC PATCH v3 3/5] perf data-convert: Add perf.data to trace.dat conversion backend Tanushree Shah @ 2026-08-03 14:59 ` Tanushree Shah 2026-08-03 15:13 ` sashiko-bot 2026-08-03 14:59 ` [RFC PATCH v3 5/5] perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat' Tanushree Shah 2026-08-06 0:18 ` [RFC PATCH v3 0/5] Add perf.data tracepoint events to trace.dat conversion Ian Rogers 5 siblings, 1 reply; 14+ messages in thread From: Tanushree Shah @ 2026-08-03 14:59 UTC (permalink / raw) To: acme, jolsa, adrian.hunter, vmolnaro, mpetlan, tmricht, maddy, irogers, namhyung Cc: linux-perf-users, linuxppc-dev, atrajeev, hbathini, Tejas.Manhas1, Tanushree.Shah, rostedt, Tanushree Shah Add new command-line option to perf data convert for generating trace.dat output files. The --to-trace-dat option: - Accepts output filename for trace.dat format - Mutually exclusive with --to-ctf and --to-json - Calls trace_convert__perf2dat() to perform conversion Usage: $ perf record -e sched:* -a sleep 1 $ perf data convert --to-trace-dat=trace.dat $ trace-cmd report trace.dat Document the new option in Documentation/perf-data.txt alongside the existing --to-ctf and --to-json entries. Signed-off-by: Tanushree Shah <tshah@linux.ibm.com> --- tools/perf/Documentation/perf-data.txt | 7 +++++ tools/perf/builtin-data.c | 43 ++++++++++++++++++++++++++ tools/perf/util/trace-dat.c | 3 ++ 3 files changed, 53 insertions(+) diff --git a/tools/perf/Documentation/perf-data.txt b/tools/perf/Documentation/perf-data.txt index 20f178d61ed7..578ba6357183 100644 --- a/tools/perf/Documentation/perf-data.txt +++ b/tools/perf/Documentation/perf-data.txt @@ -30,6 +30,13 @@ OPTIONS for 'convert' --to-json:: Triggers JSON conversion. Specify the JSON filename to output. +--to-trace-dat:: + Triggers trace.dat conversion. Converts perf.data tracepoint events + to trace.dat format v7, compatible with trace-cmd and KernelShark. + Only PERF_TYPE_TRACEPOINT events are converted. Specify the + trace.dat filename to output. Requires libtraceevent support. + Mutually exclusive with --to-ctf and --to-json. + --tod:: Convert time to wall clock time. diff --git a/tools/perf/builtin-data.c b/tools/perf/builtin-data.c index 1dd73ed6bdcb..c9c863197d02 100644 --- a/tools/perf/builtin-data.c +++ b/tools/perf/builtin-data.c @@ -30,6 +30,11 @@ static const char *data_usage[] = { static const char *to_json; static const char *to_ctf; + +#ifdef HAVE_LIBTRACEEVENT +static const char *trace_dat_output; +#endif + static struct perf_data_convert_opts opts = { .force = false, .all = false, @@ -46,6 +51,10 @@ static const struct option data_options[] = { OPT_BOOLEAN(0, "all", &opts.all, "Convert all events"), OPT_STRING(0, "time", &opts.time_str, "str", "Time span of interest (start,stop)"), +#ifdef HAVE_LIBTRACEEVENT + OPT_STRING(0, "to-trace-dat", &trace_dat_output, + "file", "Convert to trace.dat format using perf.data tracepoints"), +#endif OPT_END() }; @@ -63,10 +72,44 @@ static int cmd_data_convert(int argc, const char **argv) pr_err("You cannot specify both --to-ctf and --to-json.\n"); return -1; } +#ifdef HAVE_LIBTRACEEVENT + if (trace_dat_output && (to_json || to_ctf)) { + pr_err("You cannot specify --to-trace-dat with --to-ctf or --to-json.\n"); + return -1; + } +#endif + +#ifdef HAVE_LIBBABELTRACE_SUPPORT + #ifdef HAVE_LIBTRACEEVENT + if (!to_json && !to_ctf && !trace_dat_output) { + pr_err("You must specify one of --to-ctf, --to-json, or --to-trace-dat.\n"); + return -1; + } + #else if (!to_json && !to_ctf) { pr_err("You must specify one of --to-ctf or --to-json.\n"); return -1; } + #endif +#else + #ifdef HAVE_LIBTRACEEVENT + if (!to_json && !trace_dat_output) { + pr_err("You must specify --to-json or --to-trace-dat.\n"); + return -1; + } + #else + if (!to_json) { + pr_err("You must specify --to-json.\n"); + return -1; + } + #endif +#endif + +#ifdef HAVE_LIBTRACEEVENT + if (trace_dat_output) + return trace_convert__perf2dat(input_name ? input_name : "perf.data", + trace_dat_output, &opts); +#endif if (to_json) return bt_convert__perf2json(input_name, to_json, &opts); diff --git a/tools/perf/util/trace-dat.c b/tools/perf/util/trace-dat.c index f71e03716e27..5598496c9a61 100644 --- a/tools/perf/util/trace-dat.c +++ b/tools/perf/util/trace-dat.c @@ -861,6 +861,9 @@ void trace_dat__free_cpu_buffers(void) for (cpu = 0; cpu < trace_dat_nr_cpus; cpu++) { int i; + if (!trace_cpu_data[cpu].events) + continue; + for (i = 0; i < trace_cpu_data[cpu].count; i++) free(trace_cpu_data[cpu].events[i].raw); free(trace_cpu_data[cpu].events); -- 2.47.1 ^ permalink raw reply related [flat|nested] 14+ messages in thread
* Re: [RFC PATCH v3 4/5] perf data: Add --to-trace-dat option for converting perf.data tracepoint events into trace.dat format 2026-08-03 14:59 ` [RFC PATCH v3 4/5] perf data: Add --to-trace-dat option for converting perf.data tracepoint events into trace.dat format Tanushree Shah @ 2026-08-03 15:13 ` sashiko-bot 0 siblings, 0 replies; 14+ messages in thread From: sashiko-bot @ 2026-08-03 15:13 UTC (permalink / raw) To: Tanushree Shah; +Cc: linux-perf-users Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider: - [Medium] Memory leak of page_records and page_rec_sizes arrays when extend allocation fails. - [Medium] The --to-trace-dat conversion silently ignores the --time filtering option. - [Low] Missing explicit POSIX header inclusions in trace-dat.h and data-convert-trace.c causes build failures on musl libc. - [Low] Missing dummy inline stub for trace_convert__perf2dat when HAVE_LIBTRACEEVENT is absent. -- --- Patch [4]: [RFC PATCH v3 4/5] perf data: Add --to-trace-dat option for converting perf.data tracepoint events into trace.dat format --- commit 3f0a7b06f263871cd1c6bbd15fd1c2bd5d454bae Author: Tanushree Shah <tshah@linux.ibm.com> perf data: Add --to-trace-dat option for converting perf.data tracepoint events into trace.dat format This commit adds a new command-line option to perf data convert for generating trace.dat output files. It provides a trace_convert__perf2dat function to handle tracepoint events and write them in a format compatible with trace-cmd and KernelShark. > diff --git a/tools/perf/util/data-convert-trace.c b/tools/perf/util/data-convert-trace.c > new file mode 100644 > index 0000000000000..8dbdb2c9caa4f > --- /dev/null > +++ b/tools/perf/util/data-convert-trace.c > @@ -0,0 +1,241 @@ > +// SPDX-License-Identifier: GPL-2.0-or-later > +/* > + * Copyright 2026, IBM Corporation > + * Author: Tanushree Shah <tshah@linux.ibm.com> > + * > + * data-convert-trace.c > + * > + * Implements perf.data to trace.dat format conversion for tracepoint events. > + */ > + > +#include <errno.h> > +#include <inttypes.h> > +#include <fcntl.h> > +#include <string.h> > +#include <unistd.h> > +#include <linux/compiler.h> > +#include <linux/err.h> > + > +#include "data-convert.h" [ ... ] > +/* Convert perf.data tracepoint events to trace.dat format */ > +int trace_convert__perf2dat(const char *input, const char *to_trace, > + struct perf_data_convert_opts *opts) > +{ > + struct perf_session *session; [Severity: Medium] Does this function ignore the opts->time_str filtering option? In other format conversions, the time parameter is parsed to filter samples within the specified time ranges, but it appears to be unused here when converting to trace.dat. > + struct trace_convert tc = { > + .events_count = 0, > + }; [ ... ] > + if (!opts->force) { > + int fd = open(to_trace, O_WRONLY | O_CREAT | O_EXCL, 0644); > + > + if (fd < 0) { > + if (errno == EEXIST) > + pr_err("Output file '%s' already exists. Use --force to overwrite.\n", > + to_trace); > + else > + pr_err("Failed to open output file '%s': %s\n", > + to_trace, strerror(errno)); > + return -errno; > + } > + trace_dat_fp = fdopen(fd, "wb"); [Severity: Low] Should <stdio.h> be explicitly included at the top of this file? Standard I/O functions like fdopen() and fopen() are used here, which could cause build failures on musl libc without the direct header inclusion. > diff --git a/tools/perf/util/data-convert.h b/tools/perf/util/data-convert.h > index a96240f156712..f041c23252263 100644 > --- a/tools/perf/util/data-convert.h > +++ b/tools/perf/util/data-convert.h > @@ -19,4 +19,8 @@ int bt_convert__perf2ctf(const char *input_name, const char *to_ctf, > int bt_convert__perf2json(const char *input_name, const char *to_ctf, > struct perf_data_convert_opts *opts); > > +#ifdef HAVE_LIBTRACEEVENT > +int trace_convert__perf2dat(const char *input, const char *to_trace, > + struct perf_data_convert_opts *opts); > +#endif /* HAVE_LIBTRACEEVENT */ [Severity: Low] Is a dummy inline stub needed here for when HAVE_LIBTRACEEVENT is disabled? Without an #else block providing a fallback stub, it could cause future build failures if new callers are added that aren't guarded by the feature macro. > diff --git a/tools/perf/util/trace-dat.c b/tools/perf/util/trace-dat.c > new file mode 100644 > index 0000000000000..5598496c9a61c > --- /dev/null > +++ b/tools/perf/util/trace-dat.c [ ... ] > +static int trace_dat__write_cpu_dat(FILE *fp, struct tep_handle *pevent, > + int cpu, unsigned long long *file_offset_out) > +{ > + struct cpu_events *cpu_events = &trace_cpu_data[cpu]; [ ... ] > + for (i = 0; i < cpu_events->count; i++) { > + struct cpu_event *event = &cpu_events->events[i]; [ ... ] > + /* Emit TIME_EXTEND when delta does not fit in 27 bits */ > + if (time_delta > TRACE_DAT_RECORD_TIME_MASK) { > + unsigned int extend_hdr; > + unsigned int delta_upper; > + > + extend_size = TRACE_DAT_RECORD_TIME_EXTEND_SIZE; > + extend = calloc(1, extend_size); > + if (!extend) > + return -ENOMEM; [Severity: Medium] Can this early return cause a memory leak? If the allocation for extend fails, returning -ENOMEM directly skips the out_free label at the end of the function, which is responsible for freeing the dynamically resized page_records and page_rec_sizes arrays. > diff --git a/tools/perf/util/trace-dat.h b/tools/perf/util/trace-dat.h > new file mode 100644 > index 0000000000000..b63d4e8c1e2d7 > --- /dev/null > +++ b/tools/perf/util/trace-dat.h > @@ -0,0 +1,113 @@ > +/* SPDX-License-Identifier: GPL-2.0-or-later */ [ ... ] > +#include <stdio.h> > +#include <stdbool.h> > +#include <event-parse.h> > +#include <byteswap.h> > +#include "util.h" > + > +/* trace.dat file format version */ > +#define TRACE_DAT_VERSION "7" [ ... ] > +static inline uint16_t to_file_u16(struct tep_handle *pevent, uint16_t val) > +{ > + return tep_read_number(pevent, &val, 2); > +} [Severity: Low] Does this file need to explicitly include <stdint.h> to ensure compatibility when compiling with musl libc? Fixed-width integer types like uint16_t are used here but the standard POSIX header for them is missing. -- Sashiko AI review · https://sashiko.dev/#/patchset/20260803145958.299956-1-tshah@linux.ibm.com?part=4 ^ permalink raw reply [flat|nested] 14+ messages in thread
* [RFC PATCH v3 5/5] perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat' 2026-08-03 14:59 [RFC PATCH v3 0/5] Add perf.data tracepoint events to trace.dat conversion Tanushree Shah ` (3 preceding siblings ...) 2026-08-03 14:59 ` [RFC PATCH v3 4/5] perf data: Add --to-trace-dat option for converting perf.data tracepoint events into trace.dat format Tanushree Shah @ 2026-08-03 14:59 ` Tanushree Shah 2026-08-03 15:19 ` sashiko-bot 2026-08-06 0:18 ` [RFC PATCH v3 0/5] Add perf.data tracepoint events to trace.dat conversion Ian Rogers 5 siblings, 1 reply; 14+ messages in thread From: Tanushree Shah @ 2026-08-03 14:59 UTC (permalink / raw) To: acme, jolsa, adrian.hunter, vmolnaro, mpetlan, tmricht, maddy, irogers, namhyung Cc: linux-perf-users, linuxppc-dev, atrajeev, hbathini, Tejas.Manhas1, Tanushree.Shah, rostedt, Tanushree Shah Add a shell test covering perf data convert --to-trace-dat, alongside the existing --to-json and --to-ctf tests. The test: - skips if perf isn't linked with libtraceevent - converts a plain sched:sched_switch recording to trace.dat - converts a pipe-mode recording (perf record -o - | perf data convert -i -) - converts a recording with both tracepoint and non-tracepoint events (sched:sched_switch,cycles) - validates the resulting file with trace-cmd report if trace-cmd is installed, otherwise falls back to a basic non-empty-file check Signed-off-by: Tanushree Shah <tshah@linux.ibm.com> --- ...rf_data_converter_tracepoints_trace_dat.sh | 169 ++++++++++++++++++ 1 file changed, 169 insertions(+) create mode 100755 tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh diff --git a/tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh b/tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh new file mode 100755 index 000000000000..9ca6432618cb --- /dev/null +++ b/tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh @@ -0,0 +1,169 @@ +#!/bin/bash +# 'perf data convert --to-trace-dat' command test +# SPDX-License-Identifier: GPL-2.0-or-later +# +# Copyright 2026, IBM Corporation +# Author: Tanushree Shah <tshah@linux.ibm.com> + +set -e + +err=0 + +perfdata=$(mktemp /tmp/__perf_test.perf.data.XXXXX) +result=$(mktemp /tmp/__perf_test.output.trace.dat.XXXXX) + +cleanup() +{ + rm -f "${perfdata}" + rm -f "${result}" + trap - exit term int +} + +trap_cleanup() +{ + echo "Unexpected signal in ${FUNCNAME[1]}" + cleanup + exit 1 +} +trap trap_cleanup exit term int + +# Check if libtraceevent support is available +if ! perf check feature libtraceevent +then + echo "perf not linked with libtraceevent, skipping test" + exit 2 +fi + +# Check if trace-cmd is available for validation +have_trace_cmd=0 +if command -v trace-cmd >/dev/null 2>&1; then + have_trace_cmd=1 +fi + +check_sched_switch() +{ + if ! perf list | grep -q 'sched:sched_switch' + then + echo "sched:sched_switch tracepoint not available, skipping test" + exit 2 + fi +} + +test_trace_converter_command() +{ + echo "Testing Perf Data Conversion Command to trace.dat" + + if ! perf record -e sched:sched_switch -o "$perfdata" -- sleep 0.1 + then + echo "Failed to record perf data" + err=1 + return + fi + + rm -f "$result" + + if ! perf data convert --to-trace-dat "$result" --force -i "$perfdata" + then + echo "Perf Data Converter Command to trace.dat [FAILED]" + err=1 + return + fi + + if [ -f "$result" ] && [ -s "$result" ] ; then + echo "Perf Data Converter Command to trace.dat [SUCCESS]" + else + echo "Perf Data Converter Command to trace.dat [FAILED]" + err=1 + fi +} + +test_trace_converter_pipe() +{ + echo "Testing Perf Data Conversion Command to trace.dat (Pipe mode)" + + rm -f "$result" + + if ! perf record -e sched:sched_switch -o - -- sleep 0.1 | \ + perf data convert --to-trace-dat "$result" --force -i - + then + echo "Perf Data Converter Command to trace.dat (Pipe mode) [FAILED]" + err=1 + return + fi + + if [ -f "$result" ] && [ -s "$result" ]; then + echo "Perf Data Converter Command to trace.dat (Pipe mode) [SUCCESS]" + else + echo "Perf Data Converter Command to trace.dat (Pipe mode) [FAILED]" + err=1 + fi +} + +test_trace_converter_mixed_events() +{ + echo "Testing Perf Data Conversion with tracepoint and non-tracepoint events" + + # Record both tracepoint and non-tracepoint events + if ! perf record -e sched:sched_switch,cycles -o "$perfdata" -- sleep 0.1 + then + echo "Failed to record perf data (mixed events)" + err=1 + return + fi + + rm -f "$result" + + if ! perf data convert --to-trace-dat "$result" --force -i "$perfdata" + then + echo "Perf Data Converter with mixed events [FAILED]" + err=1 + return + fi + + if [ -f "$result" ] && [ -s "$result" ] ; then + echo "Perf Data Converter with mixed events [SUCCESS]" + else + echo "Perf Data Converter with mixed events [FAILED]" + err=1 + fi +} + +validate_trace_format() +{ + echo "Validating Perf Data Converted trace.dat file" + + if [ ! -f "$result" ] ; then + echo "File not found [FAILED]" + err=1 + return + fi + + # If trace-cmd is available, use it to validate the file + if [ "$have_trace_cmd" -eq 1 ] ; then + if trace-cmd report -i "$result" >/dev/null 2>&1 ; then + echo "trace-cmd can read the file [SUCCESS]" + else + echo "trace-cmd cannot read the file [FAILED]" + err=1 + fi + else + # Without trace-cmd, just check file exists and has content + if [ -s "$result" ] ; then + echo "Output file exists and has content [SUCCESS]" + else + echo "Output file is empty [FAILED]" + err=1 + fi + fi +} + +check_sched_switch +test_trace_converter_command +validate_trace_format +test_trace_converter_pipe +validate_trace_format +test_trace_converter_mixed_events +validate_trace_format + +cleanup +exit ${err} -- 2.47.1 ^ permalink raw reply related [flat|nested] 14+ messages in thread
* Re: [RFC PATCH v3 5/5] perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat' 2026-08-03 14:59 ` [RFC PATCH v3 5/5] perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat' Tanushree Shah @ 2026-08-03 15:19 ` sashiko-bot 2026-08-06 0:15 ` Ian Rogers 0 siblings, 1 reply; 14+ messages in thread From: sashiko-bot @ 2026-08-03 15:19 UTC (permalink / raw) To: Tanushree Shah; +Cc: linux-perf-users Thank you for your contribution! Sashiko AI review found 5 potential issue(s) to consider: - [High] Deleting a securely created temporary file and reusing its path creates a race condition vulnerable to symlink attacks. - [Medium] The exit trap incorrectly overrides intentional test skips (exit code 2), converting them into test failures (exit code 1). - [Medium] Hardcoding the 'cycles' hardware PMU event leads to brittle, spurious test failures on environments without hardware performance counters. - [Low] Stale output file verification due to incorrect cleanup ordering on error paths. - [Low] A temporary file may be leaked if a subsequent initialization step fails before the cleanup trap is registered. -- --- Patch [5]: [RFC PATCH v3 5/5] perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat' --- commit d717cd6622789cb3b2cd038490f59972dc4015dd Author: Tanushree Shah <tshah@linux.ibm.com> perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat' Add a shell test covering perf data convert --to-trace-dat, alongside the existing --to-json and --to-ctf tests. > diff --git a/tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh b/tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh > new file mode 100755 > index 0000000000000..9ca6432618cb2 > --- /dev/null > +++ b/tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh > @@ -0,0 +1,169 @@ > +#!/bin/bash [ ... ] > +perfdata=$(mktemp /tmp/__perf_test.perf.data.XXXXX) > +result=$(mktemp /tmp/__perf_test.output.trace.dat.XXXXX) [Severity: Low] Could a temporary file be leaked if an early initialization step fails? If the second mktemp fails, the script will exit immediately due to 'set -e'. Since the trap is not yet registered, the first temporary file allocated for $perfdata wouldn't be cleaned up. [ ... ] > +trap_cleanup() > +{ > + echo "Unexpected signal in ${FUNCNAME[1]}" > + cleanup > + exit 1 > +} > +trap trap_cleanup exit term int [Severity: Medium] Does this trap handler unintentionally override test skips? When the script explicitly executes 'exit 2' for a skip condition (like missing libtraceevent), the EXIT trap is still active. This handler then executes 'exit 1', which converts the skip into a spurious test failure. [ ... ] > +test_trace_converter_command() > +{ > + echo "Testing Perf Data Conversion Command to trace.dat" > + > + if ! perf record -e sched:sched_switch -o "$perfdata" -- sleep 0.1 > + then > + echo "Failed to record perf data" > + err=1 > + return > + fi > + > + rm -f "$result" [Severity: High] Does removing the securely created file path here open a symlink race condition? The script uses mktemp to securely allocate the path for $result, but deletes it before passing it to the converter with the --force flag. This leaves a time window where a predictably named path is free. If run as root, this could potentially be replaced with a symlink, allowing an arbitrary file overwrite. [ ... ] > +test_trace_converter_pipe() > +{ > + echo "Testing Perf Data Conversion Command to trace.dat (Pipe mode)" > + > + rm -f "$result" [Severity: High] This test mode appears to have the same symlink race condition as above, since the securely allocated temporary file is removed prior to use. [ ... ] > +test_trace_converter_mixed_events() > +{ > + echo "Testing Perf Data Conversion with tracepoint and non-tracepoint events" > + > + # Record both tracepoint and non-tracepoint events > + if ! perf record -e sched:sched_switch,cycles -o "$perfdata" -- sleep 0.1 [Severity: Medium] Could hardcoding the 'cycles' hardware PMU event cause spurious test failures on restricted environments? If the hardware 'cycles' event fails to open on virtual machines or unprivileged containers, perf record will return an error and fail the test step instead of skipping it or falling back to a software event. > + then > + echo "Failed to record perf data (mixed events)" > + err=1 > + return > + fi > + > + rm -f "$result" [Severity: Low] Does this cleanup ordering leave stale data available for validation on error paths? If the 'perf record' command above fails, it returns early before this file is removed. When validate_trace_format() is later called in the global flow, it will check the $result file left over from the preceding test_trace_converter_pipe() execution, which might erroneously report success for a failed command. -- Sashiko AI review · https://sashiko.dev/#/patchset/20260803145958.299956-1-tshah@linux.ibm.com?part=5 ^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [RFC PATCH v3 5/5] perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat' 2026-08-03 15:19 ` sashiko-bot @ 2026-08-06 0:15 ` Ian Rogers 2026-08-06 9:20 ` Tanushree Shah 0 siblings, 1 reply; 14+ messages in thread From: Ian Rogers @ 2026-08-06 0:15 UTC (permalink / raw) To: sashiko-reviews; +Cc: Tanushree Shah, linux-perf-users On Mon, Aug 3, 2026 at 9:50 AM <sashiko-bot@kernel.org> wrote: > > Thank you for your contribution! Sashiko AI review found 5 potential issue(s) to consider: > - [High] Deleting a securely created temporary file and reusing its path creates a race condition vulnerable to symlink attacks. > - [Medium] The exit trap incorrectly overrides intentional test skips (exit code 2), converting them into test failures (exit code 1). > - [Medium] Hardcoding the 'cycles' hardware PMU event leads to brittle, spurious test failures on environments without hardware performance counters. > - [Low] Stale output file verification due to incorrect cleanup ordering on error paths. > - [Low] A temporary file may be leaked if a subsequent initialization step fails before the cleanup trap is registered. > -- > > --- Patch [5]: [RFC PATCH v3 5/5] perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat' --- > commit d717cd6622789cb3b2cd038490f59972dc4015dd > Author: Tanushree Shah <tshah@linux.ibm.com> > > perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat' > > Add a shell test covering perf data convert --to-trace-dat, > alongside the existing --to-json and --to-ctf tests. > > > diff --git a/tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh b/tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh > > new file mode 100755 > > index 0000000000000..9ca6432618cb2 > > --- /dev/null > > +++ b/tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh > > @@ -0,0 +1,169 @@ > > +#!/bin/bash > > [ ... ] > > > +perfdata=$(mktemp /tmp/__perf_test.perf.data.XXXXX) > > +result=$(mktemp /tmp/__perf_test.output.trace.dat.XXXXX) > > [Severity: Low] > Could a temporary file be leaked if an early initialization step fails? > > If the second mktemp fails, the script will exit immediately due to 'set -e'. > Since the trap is not yet registered, the first temporary file allocated for > $perfdata wouldn't be cleaned up. > > [ ... ] > > > +trap_cleanup() > > +{ > > + echo "Unexpected signal in ${FUNCNAME[1]}" > > + cleanup > > + exit 1 > > +} > > +trap trap_cleanup exit term int > > [Severity: Medium] > Does this trap handler unintentionally override test skips? > > When the script explicitly executes 'exit 2' for a skip condition (like > missing libtraceevent), the EXIT trap is still active. This handler then > executes 'exit 1', which converts the skip into a spurious test failure. > > [ ... ] > > > +test_trace_converter_command() > > +{ > > + echo "Testing Perf Data Conversion Command to trace.dat" > > + > > + if ! perf record -e sched:sched_switch -o "$perfdata" -- sleep 0.1 > > + then > > + echo "Failed to record perf data" > > + err=1 > > + return > > + fi > > + > > + rm -f "$result" > > [Severity: High] > Does removing the securely created file path here open a symlink race > condition? > > The script uses mktemp to securely allocate the path for $result, but deletes > it before passing it to the converter with the --force flag. This leaves > a time window where a predictably named path is free. If run as root, this > could potentially be replaced with a symlink, allowing an arbitrary file > overwrite. > > [ ... ] > > > +test_trace_converter_pipe() > > +{ > > + echo "Testing Perf Data Conversion Command to trace.dat (Pipe mode)" > > + > > + rm -f "$result" > > [Severity: High] > This test mode appears to have the same symlink race condition as above, > since the securely allocated temporary file is removed prior to use. > > [ ... ] > > > +test_trace_converter_mixed_events() > > +{ > > + echo "Testing Perf Data Conversion with tracepoint and non-tracepoint events" > > + > > + # Record both tracepoint and non-tracepoint events > > + if ! perf record -e sched:sched_switch,cycles -o "$perfdata" -- sleep 0.1 > > [Severity: Medium] > Could hardcoding the 'cycles' hardware PMU event cause spurious test failures > on restricted environments? > > If the hardware 'cycles' event fails to open on virtual machines or > unprivileged containers, perf record will return an error and fail the test > step instead of skipping it or falling back to a software event. As the `perf record` lacks the `-a` flag this will just record events for the sleep command and the number of cycles there will be limited. I'd suggest making this: perf record -e sched:sched_switch,cpu-clock -a -o "$perfdata" -- sleep 0.1 The cpu-clock is a perf event based on a timer interrupt and should always exist whether the system be an S390 or a limited system such as running in a hypervisor. Thanks, Ian > > + then > > + echo "Failed to record perf data (mixed events)" > > + err=1 > > + return > > + fi > > + > > + rm -f "$result" > > [Severity: Low] > Does this cleanup ordering leave stale data available for validation on error > paths? > > If the 'perf record' command above fails, it returns early before this file > is removed. When validate_trace_format() is later called in the global flow, > it will check the $result file left over from the preceding > test_trace_converter_pipe() execution, which might erroneously report success > for a failed command. > > -- > Sashiko AI review · https://sashiko.dev/#/patchset/20260803145958.299956-1-tshah@linux.ibm.com?part=5 > ^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [RFC PATCH v3 5/5] perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat' 2026-08-06 0:15 ` Ian Rogers @ 2026-08-06 9:20 ` Tanushree Shah 0 siblings, 0 replies; 14+ messages in thread From: Tanushree Shah @ 2026-08-06 9:20 UTC (permalink / raw) To: Ian Rogers, sashiko-reviews; +Cc: linux-perf-users On 06/08/26 05:45, Ian Rogers wrote: > On Mon, Aug 3, 2026 at 9:50 AM <sashiko-bot@kernel.org> wrote: >> >> Thank you for your contribution! Sashiko AI review found 5 potential issue(s) to consider: >> - [High] Deleting a securely created temporary file and reusing its path creates a race condition vulnerable to symlink attacks. >> - [Medium] The exit trap incorrectly overrides intentional test skips (exit code 2), converting them into test failures (exit code 1). >> - [Medium] Hardcoding the 'cycles' hardware PMU event leads to brittle, spurious test failures on environments without hardware performance counters. >> - [Low] Stale output file verification due to incorrect cleanup ordering on error paths. >> - [Low] A temporary file may be leaked if a subsequent initialization step fails before the cleanup trap is registered. >> -- >> >> --- Patch [5]: [RFC PATCH v3 5/5] perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat' --- >> commit d717cd6622789cb3b2cd038490f59972dc4015dd >> Author: Tanushree Shah <tshah@linux.ibm.com> >> >> perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat' >> >> Add a shell test covering perf data convert --to-trace-dat, >> alongside the existing --to-json and --to-ctf tests. >> >>> diff --git a/tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh b/tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh >>> new file mode 100755 >>> index 0000000000000..9ca6432618cb2 >>> --- /dev/null >>> +++ b/tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh >>> @@ -0,0 +1,169 @@ >>> +#!/bin/bash >> >> [ ... ] >> >>> +perfdata=$(mktemp /tmp/__perf_test.perf.data.XXXXX) >>> +result=$(mktemp /tmp/__perf_test.output.trace.dat.XXXXX) >> >> [Severity: Low] >> Could a temporary file be leaked if an early initialization step fails? >> >> If the second mktemp fails, the script will exit immediately due to 'set -e'. >> Since the trap is not yet registered, the first temporary file allocated for >> $perfdata wouldn't be cleaned up. >> >> [ ... ] >> >>> +trap_cleanup() >>> +{ >>> + echo "Unexpected signal in ${FUNCNAME[1]}" >>> + cleanup >>> + exit 1 >>> +} >>> +trap trap_cleanup exit term int >> >> [Severity: Medium] >> Does this trap handler unintentionally override test skips? >> >> When the script explicitly executes 'exit 2' for a skip condition (like >> missing libtraceevent), the EXIT trap is still active. This handler then >> executes 'exit 1', which converts the skip into a spurious test failure. >> >> [ ... ] >> >>> +test_trace_converter_command() >>> +{ >>> + echo "Testing Perf Data Conversion Command to trace.dat" >>> + >>> + if ! perf record -e sched:sched_switch -o "$perfdata" -- sleep 0.1 >>> + then >>> + echo "Failed to record perf data" >>> + err=1 >>> + return >>> + fi >>> + >>> + rm -f "$result" >> >> [Severity: High] >> Does removing the securely created file path here open a symlink race >> condition? >> >> The script uses mktemp to securely allocate the path for $result, but deletes >> it before passing it to the converter with the --force flag. This leaves >> a time window where a predictably named path is free. If run as root, this >> could potentially be replaced with a symlink, allowing an arbitrary file >> overwrite. >> >> [ ... ] >> >>> +test_trace_converter_pipe() >>> +{ >>> + echo "Testing Perf Data Conversion Command to trace.dat (Pipe mode)" >>> + >>> + rm -f "$result" >> >> [Severity: High] >> This test mode appears to have the same symlink race condition as above, >> since the securely allocated temporary file is removed prior to use. >> >> [ ... ] >> >>> +test_trace_converter_mixed_events() >>> +{ >>> + echo "Testing Perf Data Conversion with tracepoint and non-tracepoint events" >>> + >>> + # Record both tracepoint and non-tracepoint events >>> + if ! perf record -e sched:sched_switch,cycles -o "$perfdata" -- sleep 0.1 >> >> [Severity: Medium] >> Could hardcoding the 'cycles' hardware PMU event cause spurious test failures >> on restricted environments? >> >> If the hardware 'cycles' event fails to open on virtual machines or >> unprivileged containers, perf record will return an error and fail the test >> step instead of skipping it or falling back to a software event. > > As the `perf record` lacks the `-a` flag this will just record events > for the sleep command and the number of cycles there will be limited. > I'd suggest making this: > perf record -e sched:sched_switch,cpu-clock -a -o "$perfdata" -- sleep 0.1 > The cpu-clock is a perf event based on a timer interrupt and should > always exist whether the system be an S390 or a limited system such as > running in a hypervisor. > > Thanks, > Ian > Thanks Ian for the suggestion, It makes sense. I will change the command to "perf record -e sched:sched_switch,cpu-clock -a -o "$perfdata" -- sleep 0.1". Thanks Tanushree Shah >>> + then >>> + echo "Failed to record perf data (mixed events)" >>> + err=1 >>> + return >>> + fi >>> + >>> + rm -f "$result" >> >> [Severity: Low] >> Does this cleanup ordering leave stale data available for validation on error >> paths? >> >> If the 'perf record' command above fails, it returns early before this file >> is removed. When validate_trace_format() is later called in the global flow, >> it will check the $result file left over from the preceding >> test_trace_converter_pipe() execution, which might erroneously report success >> for a failed command. >> >> -- >> Sashiko AI review · https://sashiko.dev/#/patchset/20260803145958.299956-1-tshah@linux.ibm.com?part=5 >> ^ permalink raw reply [flat|nested] 14+ messages in thread
* Re: [RFC PATCH v3 0/5] Add perf.data tracepoint events to trace.dat conversion 2026-08-03 14:59 [RFC PATCH v3 0/5] Add perf.data tracepoint events to trace.dat conversion Tanushree Shah ` (4 preceding siblings ...) 2026-08-03 14:59 ` [RFC PATCH v3 5/5] perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat' Tanushree Shah @ 2026-08-06 0:18 ` Ian Rogers 5 siblings, 0 replies; 14+ messages in thread From: Ian Rogers @ 2026-08-06 0:18 UTC (permalink / raw) To: Tanushree Shah Cc: acme, jolsa, adrian.hunter, vmolnaro, mpetlan, tmricht, maddy, namhyung, linux-perf-users, linuxppc-dev, atrajeev, hbathini, Tejas.Manhas1, Tanushree.Shah, rostedt On Mon, Aug 3, 2026 at 8:00 AM Tanushree Shah <tshah@linux.ibm.com> wrote: > > This RFC patch series introduces support for converting perf.data files > containing tracepoint events into trace.dat format, enabling seamless > visualization and analysis using KernelShark. > > ====================== > Background and Motivation > ====================== > > Currently, perf and trace-cmd operate as separate tracing ecosystems with > incompatible data formats. Users who collect tracepoint data with > 'perf record' cannot easily visualize it in KernelShark's graphical > timeline view or leverage trace-cmd's analysis capabilities. > > This creates workflow friction when users need to: > > - Visualize perf tracepoint data in KernelShark's interactive graphical > timeline > - Share trace data between perf and trace-cmd workflows and toolchains > - Perform architecture-independent conversion and analysis of traces > > This conversion bridge eliminates these barriers by enabling seamless > data exchange between perf and trace-cmd ecosystems, allowing users to > choose the best tool for each analysis phase. > > ====================== > Implementation Overview > ====================== > > The series implements the trace.dat file format specification (version 7) > within perf's data conversion framework. > > **Patch 1/5: Core trace.dat Export Infrastructure** > Introduces util/trace-dat.c and util/trace-dat.h implementing: > - Per-CPU raw event buffer management (init, collect, free) > - Ftrace ring buffer page construction > - trace.dat section writers (strings, options, flyrecord sections) > > **Patch 2/5: Metadata Integration** > Extends util/trace-event-read.c to write trace.dat metadata during > perf.data > parsing: > - Initial format header (magic, version, endian, page size, compression) > - Section 16: HEADER INFO (header_page + header_event) > - Section 17: FTRACE EVENT FORMATS > - Section 18: EVENT FORMATS (per system/event format files) > - Section 19: KALLSYMS > - Section 21: CMDLINES > - Section 15: STRINGS (written last after all sections) > > **Patch 3/5: Conversion Backend** > Implements util/data-convert-trace.c with trace_convert__perf2dat() > function: > - Processes PERF_TYPE_TRACEPOINT samples via process_sample_event() > - Collects raw event data per-CPU using trace_dat__collect_cpu_event() > - Writes OPTIONS sections (CPUCOUNT, TRACECLOCK, metadata offsets) > - Writes FLYRECORD section with per-CPU ring buffer pages > > **Patch 4/5: User Interface** > Extends tools/perf/builtin-data.c with --to-trace-dat option: > - Adds command-line option for trace.dat output > - Mutually exclusive with --to-ctf and --to-json > - Calls trace_convert__perf2dat() to perform conversion > > **Patch 5/5: Shell Test** > Adds a shell test (tools/perf/tests/shell/) covering normal tracepoint > recordings, pipe mode and mixed tracepoint/non-tracepoint recording > conversions and --force flag behaviour. > > ====================== > Current Implementation Details > ====================== > > **trace.dat Format Version:** > The implementation currently targets trace.dat format version 7, which > is the stable version supported by current trace-cmd releases (v3.x). > This version is hardcoded to ensure compatibility with existing > trace-cmd and KernelShark installations. Future enhancements could add > version negotiation or support for newer format versions as they become > standardized. > > **Compression Strategy:** > Compression is explicitly disabled (set to NONE) in the generated > trace.dat files. > This design choice: > - Simplifies the initial implementation and testing > - Ensures maximum compatibility across trace-cmd versions > - Avoids external compression library dependencies > > Future work could add support for various compression algorithms (zlib, > zstd, lz4) with runtime selection via command-line options, significantly > reducing file sizes for large traces. > > ====================== > Usage Example > ====================== > > ```bash > *Record tracepoint events with perf* > perf record -e sched:sched_switch -e sched:sched_wakeup -a sleep 10 > > *Convert to trace.dat format* > perf data convert --to-trace-dat=output.dat > > *Verify trace.dat structure* > trace-cmd dump --summary output.dat > > *Analyze with trace-cmd* > trace-cmd report output.dat > > *Visualize in KernelShark* > kernelshark output.dat > ``` > > **Conversion Output:** > ``` > [ perf data convert: Converted 'perf.data' into trace.dat format > 'output.dat' ] > [ perf data convert: Converted 2684 events ] > ``` > **trace-cmd dump --summary Output:** > ``` > Tracing meta data in file output.dat: > [Initial format] > 7 [Version] > 0 [Little endian] > 8 [Bytes in a long] > 65536 [Page size, bytes] > none [Compression algorithm] > [Compression version] > [buffer "", "local" clock, 65536 page size, 16 cpus, 1048576 bytes > flyrecord data] > [10 options] > [Saved command lines, 0 bytes] > [Kallsyms, 0 bytes] > [Ftrace format, 0 events] > [Header page, 206 bytes] > [Header event, 205 bytes] > [Events format, 1 systems] > [9 sections] > ``` > ====================== > Testing and Verification > ====================== > > The series has been extensively tested with: > - Various tracepoint events (sched, irq, syscalls, block I/O) > - Mixed recordings containing both tracepoint and non-tracepoint events > only tracepoints converted) > - Verification with trace-cmd report and KernelShark visualization > - Memory leak testing with Valgrind (0 bytes leaked). > - Cross-architecture testing: v1 tested x86_64 and ppc64le. v2 adds > s390 (big-endian) perf.data converted on both ppc64le > (little-endian) and x86_64 (little-endian) hosts, in addition to > same-arch x86_64 (LE->LE) and ppc64le (BE->BE) conversion. > - Pipe mode support has been tested end-to-end. (in v2) > > All generated trace.dat files successfully open in: > - trace-cmd report (v3.1+) > - KernelShark (v2.0+) > > > ====================== > Next Steps > ====================== > > We would highly appreciate reviews, comments, and feedback on: > - The overall architectural approach and integration points > - Compatibility considerations with trace-cmd ecosystem > - Performance characteristics for large-scale traces > - Additional use cases or workflow scenarios > - Future enhancement priorities > > --- > Changes in v3 > > - Rebase on latest perf-tools-next and resolve merge conflicts. > - Drop evsel parameter from process_sample_event() following upstream > commit "perf tool: Remove evsel from tool APIs that pass the sample". > > Changes in v2 > > Addressing the Sashiko AI review findings on v1: > > Cross-arch correctness: > - Introduce to_file_u16/u32/u64 helpers (wrapping tep_read_number()) > to write all multi-byte fields in the recorded machine's byte order; > apply throughout metadata sections and flyrecord page/record headers > (ts, commit, TIME_EXTEND, large-event data_len). > - Fix flyrecord record header bit layout for big-endian files. > The record header word bit layout differs by file endianness, > matching kbuffer-parse.c type_len4host()/ts4host(): > LE: type_len in bits [4:0], time_delta in bits [31:5] > BE: type_len in bits [31:27], time_delta in bits [26:0] > > Pipe mode: > - Add process_attr(), process_feature(), and process_tracing_data() > callbacks required for pipe mode operation. > - Defer CPU buffer initialisation until the first tracepoint sample, > after process_feature()/process_tracing_data() have populated the > session header. This ensures the recorded machine's CPU count is > used rather than the host's - critical for cross-platform analysis. > > Format compliance: > - Implement TIME_EXTEND records for timestamp deltas >27 bits to > prevent silent truncation and maintain chronological ordering. > - Fix large event encoding (>=29 words): use type_len=0 with a > separate 32-bit length word, avoiding collision with reserved types > (PADDING=29, TIME_EXTEND=30, TIME_STAMP=31). > - Add bounds check rejecting records larger than a page payload before > batching, preventing heap overflow in trace_dat__write_page(). > - Fix flyrecord section_size to exclude the 16-byte section header, > matching trace.dat specification and trace-cmd behaviour. > > CLI behavior: > - Fix --force flag: open with O_CREAT|O_EXCL when force is not set, > failing with -EEXIST instead of silently overwriting existing files. > > Memory safety: > - Fix realloc overwrite of cpu_events->events and page_records on > failure: use temporary pointers, only commit on success. > - Fix use-after-free/double-free in sequential page_records realloc > failure: replace with malloc+memcpy+free pattern. > - Fix section_size computed from before section header position. > - Add NULL checks for get_tracing_file(), calloc() padding, and > trace_dat_options_offset assignment on write failure. > - Use goto out_free on record allocation failure to avoid leaking > accumulated page_records entries. > - Replace direct read() with do_read() in read_proc_kallsyms() to > handle short reads correctly. > - On fwrite failure, set trace_dat_write_failed and continue parsing > so that perf.data processing completes normally. > > Testing (new in v2): > - Add shell test covering conversion, trace-cmd dump validation, > sched_switch event verification, and --force flag behaviour. > > Documentation (new in v2): > - Add documentation for 'perf data convert --to-trace-dat', covering > usage and supported options. > > v1: https://lore.kernel.org/linux-perf-users/20260608125951.90425-2-tshah@linux.ibm.com/ Still looks great, and it's getting better as you address the nits Sashiko is raising. I look forward to testing. Thanks, Ian > Tanushree Shah (5): > perf/trace-dat: Add trace.dat export infrastructure > perf/trace-event: Write trace.dat metadata sections during parsing > perf data-convert: Add perf.data to trace.dat conversion backend > perf data: Add --to-trace-dat option for converting perf.data > tracepoint events into trace.dat format > perf test: Add test validating trace.dat generated by 'perf data > convert --to-trace-dat' > > tools/perf/Documentation/perf-data.txt | 7 + > tools/perf/builtin-data.c | 43 + > ...rf_data_converter_tracepoints_trace_dat.sh | 169 ++++ > tools/perf/util/Build | 2 + > tools/perf/util/data-convert-trace.c | 241 +++++ > tools/perf/util/data-convert.h | 4 + > tools/perf/util/trace-dat.c | 876 ++++++++++++++++++ > tools/perf/util/trace-dat.h | 113 +++ > tools/perf/util/trace-event-read.c | 307 +++++- > 9 files changed, 1755 insertions(+), 7 deletions(-) > create mode 100755 tools/perf/tests/shell/test_perf_data_converter_tracepoints_trace_dat.sh > create mode 100644 tools/perf/util/data-convert-trace.c > create mode 100644 tools/perf/util/trace-dat.c > create mode 100644 tools/perf/util/trace-dat.h > > -- > 2.47.1 > ^ permalink raw reply [flat|nested] 14+ messages in thread
end of thread, other threads:[~2026-08-06 9:20 UTC | newest] Thread overview: 14+ messages (download: mbox.gz follow: Atom feed -- links below jump to the message on this page -- 2026-08-03 14:59 [RFC PATCH v3 0/5] Add perf.data tracepoint events to trace.dat conversion Tanushree Shah 2026-08-03 14:59 ` [RFC PATCH v3 1/5] perf/trace-dat: Add trace.dat export infrastructure Tanushree Shah 2026-08-03 15:14 ` sashiko-bot 2026-08-03 14:59 ` [RFC PATCH v3 2/5] perf/trace-event: Write trace.dat metadata sections during parsing Tanushree Shah 2026-08-03 15:14 ` sashiko-bot 2026-08-03 14:59 ` [RFC PATCH v3 3/5] perf data-convert: Add perf.data to trace.dat conversion backend Tanushree Shah 2026-08-03 15:11 ` sashiko-bot 2026-08-03 14:59 ` [RFC PATCH v3 4/5] perf data: Add --to-trace-dat option for converting perf.data tracepoint events into trace.dat format Tanushree Shah 2026-08-03 15:13 ` sashiko-bot 2026-08-03 14:59 ` [RFC PATCH v3 5/5] perf test: Add test validating trace.dat generated by 'perf data convert --to-trace-dat' Tanushree Shah 2026-08-03 15:19 ` sashiko-bot 2026-08-06 0:15 ` Ian Rogers 2026-08-06 9:20 ` Tanushree Shah 2026-08-06 0:18 ` [RFC PATCH v3 0/5] Add perf.data tracepoint events to trace.dat conversion Ian Rogers
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox