intel-gfx.lists.freedesktop.org archive mirror
 help / color / mirror / Atom feed
From: Lionel Landwerlin <lionel.g.landwerlin@intel.com>
To: Sagar Arun Kamble <sagar.a.kamble@intel.com>,
	Robert Bragg <robert@sixbynine.org>
Cc: Intel Graphics Development <intel-gfx@lists.freedesktop.org>
Subject: Re: [RFC 0/4] GPU/CPU timestamps correlation for relating OA samples with system events
Date: Thu, 28 Dec 2017 17:13:01 +0000	[thread overview]
Message-ID: <ed173123-6d6b-9231-bbbc-4d5094c42c57@intel.com> (raw)
In-Reply-To: <04eca028-3705-5a28-b500-089ca19e712c@intel.com>


[-- Attachment #1.1: Type: text/plain, Size: 5395 bytes --]

On 26/12/17 05:32, Sagar Arun Kamble wrote:
>
>
>
> On 12/22/2017 3:46 PM, Lionel Landwerlin wrote:
>> On 22/12/17 09:30, Sagar Arun Kamble wrote:
>>>
>>>
>>>
>>> On 12/21/2017 6:29 PM, Lionel Landwerlin wrote:
>>>> Some more findings I made while playing with this series & GPUTop.
>>>> Turns out the 2ms drift per second is due to timecounter. Adding 
>>>> the delta this way :
>>>>
>>>> https://github.com/djdeath/linux/commit/7b002cb360483e331053aec0f98433a5bd5c5c3f#diff-9b74bd0cfaa90b601d80713c7bd56be4R607
>>>>
>>>> Eliminates the drift.
>>> I see two imp. changes 1. approximation of start time during 
>>> init_timecounter 2. overflow handling in delta accumulation.
>>> With these incorporated, I guess timecounter should also work in 
>>> same fashion.
>>
>> I think the arithmetic in timecounter is inherently lossy and that's 
>> why we're seeing a drift.
> Could you share details about platform, scenario in which 2ms drift 
> per second is being seen with timecounter.
> I did not observe this on SKL.

The 2ms drift was on SKL GT4.

With the patch above, I'm seeing only a ~40us drift over ~7seconds of 
recording both perf tracepoints & i915 perf reports.
I'm tracking the kernel tracepoints adding gem requests and the i915 
perf reports.
Here a screenshot at the beginning of the 7s recording : 
https://i.imgur.com/hnexgjQ.png (you can see the gem request add before 
the work starts in the i915 perf reports).
At the end of the recording, the gem requests appear later than the work 
in the i915 perf report : https://i.imgur.com/oCd0C9T.png

I'll try to prepare some IGT tests that show the drift using perf & i915 
perf, so we can run those on different platforms.
I tend to mostly test on a SKL GT4 & KBL GT2, but BXT definitely needs 
more attention...

>> Could we be using it wrong?
>>
> if we use two changes highlighted above with timecounter maybe we will 
> get same results as your current implementation.
>> In the patch above, I think there is still a drift because of the 
>> potential fractional part loss at every delta we add.
>> But it should only be a fraction of a nanosecond multiplied by the 
>> number of reports over a period of time.
>> With a report every 1us, that should still be much less than a 1ms of 
>> drift over 1s.
>>
> timecounter interface takes care of fractional parts so that should 
> help us.
> we can either go with timecounter or our own implementation provided 
> conversions are precise.

Looking at clocks_calc_mult_shift(), it seems clear to me that there is 
less precision when using timecounter :

  /*
   * Find the conversion shift/mult pair which has the best
   * accuracy and fits the maxsec conversion range:
   */

On the other hand, there is a performance penalty for doing a div64 for 
every report.

>> We can probably do better by always computing the clock using the 
>> entire delta rather than the accumulated delta.
>>
> issue is that the reported clock cycles in the OA report is 32bits LSB 
> of GPU TS whereas counter is 36bits. Hence we will need to
> accumulate the delta. ofc there is assumption that two reports can't 
> be spaced with count value of 0xffffffff apart.

You're right :)
I thought maybe we could do this :

Look at teduhe opening period parameter, if it's superior to the period 
of timestamps wrapping, make sure we schle some work on kernel context 
to generate a context switch report (like at least once every 6 minutes 
on gen9).

>>>> Timelines of perf i915 tracepoints & OA reports now make a lot more 
>>>> sense.
>>>>
>>>> There is still the issue that reading the CPU clock & the RCS 
>>>> timestamp is inherently not atomic. So there is a delta there.
>>>> I think we should add a new i915 perf record type to express the 
>>>> delta that we measure this way :
>>>>
>>>> https://github.com/djdeath/linux/commit/7b002cb360483e331053aec0f98433a5bd5c5c3f#diff-9b74bd0cfaa90b601d80713c7bd56be4R2475
>>>>
>>>> So that userspace knows there might be a global offset between the 
>>>> 2 times and is able to present it.
>>> agree on this. Delta ns1-ns0 can be interpreted as max drift.
>>>> Measurement on my KBL system were in the order of a few 
>>>> microseconds (~30us).
>>>> I guess we might be able to setup the correlation point better 
>>>> (masking interruption?) to reduce the delta.
>>> already using spin_lock. Do you mean NMI?
>>
>> I don't actually know much on this point.
>> if spin_lock is the best we can do, then that's it :)
>>
>>>>
>>>> Thanks,
>>>>
>>>> -
>>>> Lionel
>>>>
>>>>
>>>> On 07/12/17 00:57, Robert Bragg wrote:
>>>>>
>>>>>
>>>>> On Thu, Dec 7, 2017 at 12:48 AM, Robert Bragg 
>>>>> <robert@sixbynine.org <mailto:robert@sixbynine.org>> wrote:
>>>>>
>>>>>
>>>>>     at least from what I wrote back then it looks like I was
>>>>>     seeing a drift of a few milliseconds per second on SKL. I
>>>>>     vaguely recall it being much worse given the frequency
>>>>>     constants we had for Haswell.
>>>>>
>>>>>
>>>>> Sorry I didn't actually re-read my own message properly before 
>>>>> referencing it :) Apparently the 2ms per second drift was for 
>>>>> Haswell, so presumably not quite so bad for SKL.
>>>>>
>>>>> - Robert
>>>>>
>>>>>
>>>>>
>>>>> _______________________________________________
>>>>> Intel-gfx mailing list
>>>>> Intel-gfx@lists.freedesktop.org
>>>>> https://lists.freedesktop.org/mailman/listinfo/intel-gfx
>>>>
>>>>
>>>
>>
>


[-- Attachment #1.2: Type: text/html, Size: 11534 bytes --]

[-- Attachment #2: Type: text/plain, Size: 160 bytes --]

_______________________________________________
Intel-gfx mailing list
Intel-gfx@lists.freedesktop.org
https://lists.freedesktop.org/mailman/listinfo/intel-gfx

  reply	other threads:[~2017-12-28 17:13 UTC|newest]

Thread overview: 35+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2017-11-15 12:13 [RFC 0/4] GPU/CPU timestamps correlation for relating OA samples with system events Sagar Arun Kamble
2017-11-15 12:13 ` [RFC 1/4] drm/i915/perf: Add support to correlate GPU timestamp with system time Sagar Arun Kamble
2017-11-15 12:25   ` Chris Wilson
2017-11-15 16:41     ` Sagar Arun Kamble
2017-12-05 13:58     ` Lionel Landwerlin
2017-12-06  8:17       ` Sagar Arun Kamble
2017-11-15 12:13 ` [RFC 2/4] drm/i915/perf: Add support for collecting 64 bit timestamps with OA reports Sagar Arun Kamble
2017-12-06 16:01   ` Lionel Landwerlin
2017-12-21  8:38     ` Sagar Arun Kamble
2017-11-15 12:13 ` [RFC 3/4] drm/i915/perf: Extract raw GPU timestamps from " Sagar Arun Kamble
2017-12-06 19:55   ` Lionel Landwerlin
2017-12-21  8:50     ` Sagar Arun Kamble
2017-11-15 12:13 ` [RFC 4/4] drm/i915/perf: Send system clock monotonic time in perf samples Sagar Arun Kamble
2017-11-15 12:31   ` Chris Wilson
2017-11-15 16:51     ` Sagar Arun Kamble
2017-11-15 17:54   ` Sagar Arun Kamble
2017-12-05 14:22   ` Lionel Landwerlin
2017-12-06  8:31     ` Sagar Arun Kamble
2017-11-15 12:30 ` ✗ Fi.CI.BAT: warning for GPU/CPU timestamps correlation for relating OA samples with system events Patchwork
2017-12-05 14:16 ` [RFC 0/4] " Lionel Landwerlin
2017-12-05 14:28   ` Robert Bragg
2017-12-05 14:37     ` Lionel Landwerlin
2017-12-06  9:01       ` Sagar Arun Kamble
2017-12-06 20:02 ` Lionel Landwerlin
2017-12-22  5:15   ` Sagar Arun Kamble
2017-12-22  5:26     ` Sagar Arun Kamble
2017-12-07  0:48 ` Robert Bragg
2017-12-07  0:57   ` Robert Bragg
2017-12-21 12:59     ` Lionel Landwerlin
2017-12-22  9:30       ` Sagar Arun Kamble
2017-12-22 10:16         ` Lionel Landwerlin
2017-12-26  5:32           ` Sagar Arun Kamble
2017-12-28 17:13             ` Lionel Landwerlin [this message]
2018-01-03  5:38               ` Sagar Arun Kamble
2017-12-22  6:06   ` Sagar Arun Kamble

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=ed173123-6d6b-9231-bbbc-4d5094c42c57@intel.com \
    --to=lionel.g.landwerlin@intel.com \
    --cc=intel-gfx@lists.freedesktop.org \
    --cc=robert@sixbynine.org \
    --cc=sagar.a.kamble@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for NNTP newsgroup(s).