]> git.hungrycats.org Git - linux/log
linux
3 months agodrm/amd/display: Add KUnit tests for amdgpu_dm_quirks
Alex Hung [Thu, 7 May 2026 16:56:40 +0000 (10:56 -0600)]
drm/amd/display: Add KUnit tests for amdgpu_dm_quirks

Add KUnit test file amdgpu_dm_quirks_test.c covering retrieve_dmi_info().

Three test cases are provided:
- Verify aux_hpd_discon_quirk is reset to false even when previously true
- Verify edp0_on_dp1_quirk is reset to false even when previously true
- Verify both quirks remain false on a zero-initialised dm when no
  DMI match is found (expected in UML/KUnit environment)

Register the new test object in the tests/Makefile under
CONFIG_DRM_AMD_DC_KUNIT_TEST.

Assisted-by: Copilot:Claude-Sonnet-4.6
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Add KUnit tests for amdgpu_dm_helpers
Alex Hung [Wed, 6 May 2026 22:39:07 +0000 (16:39 -0600)]
drm/amd/display: Add KUnit tests for amdgpu_dm_helpers

Add amdgpu_dm_helpers_test.c with 32 KUnit test cases covering the
following functions in amdgpu_dm_helpers.c:

- edid_extract_panel_id(): basic extraction with known mfg_id and
  prod_code; zero inputs produce zero output.
- dm_is_freesync_pcon_whitelist(): every entry in the whitelist
  table returns true; an unknown ID and a zero ID return false.
- populate_hdmi_info_from_connector(): scdc_present is copied from
  hdmi->scdc.supported for both true and false; FRL DSC fields map
  10bpc and 12bpc correctly and ignore unknown values.
- dm_get_adaptive_sync_support_type(): five cases covering the
  default non-converter path, HDMI converter without conditions,
  partial conditions, all conditions met with a whitelist device
  (FREESYNC_TYPE_PCON_IN_WHITELIST), and all conditions met with a
  non-whitelisted device.
- dm_helpers_is_fullscreen() / dm_helpers_is_hdr_on(): stubs always
  return false.
- get_max_frl_rate(): all six valid lane/rate combinations plus the
  unknown combination returning 0.
- dm_dtn_log_begin()/dm_dtn_log_append_v()/dm_dtn_log_end(): buffer
  accumulation and NULL-context handling without crashing.
- dm_helpers_dp_read_dpcd()/dm_helpers_dp_write_dpcd(): NULL link
  private data returns false.
- dm_helpers_dp_mst_start_top_mgr()/dm_helpers_dp_mst_stop_top_mgr():
  NULL link private data and the boot path.
- dm_helpers_dp_write_hblank_reduction(): stub returns false.

Assisted-by: Copilot:Claude-Opus-4.8
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Add KUnit tests for amdgpu_dm_services
Alex Hung [Wed, 6 May 2026 21:54:47 +0000 (15:54 -0600)]
drm/amd/display: Add KUnit tests for amdgpu_dm_services

Add amdgpu_dm_services_test.c with KUnit coverage for five
functions in amdgpu_dm_services.c:

- dm_get_elapse_time_in_ns(): four arithmetic cases covering
  zero delta, positive delta, ULLONG_MAX span, and unsigned
  wraparound.
- dm_perf_trace_timestamp(): one case verifying the function
  dereferences ctx->perf_trace safely (the tracepoint is a
  no-op without an attached probe).
- dm_trace_smu_enter(): two cases for the empty stub with NULL
  ctx and with non-zero parameters.
- dm_trace_smu_exit(): three cases for the empty stub covering
  success, failure, and a non-zero response value.
- dm_query_extended_brightness_caps(): four guard-clause cases
  (NULL ctx, NULL caps, NULL ctx->driver_context, NULL ctx with
  LCD2) plus two success cases covering the LCD1 slot with
  luminance data copy and a non-LCD1 display using the second
  backlight slot with zero data points.

Assisted-by: Copilot:Claude-Opus-4.8
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Add KUnit tests for amdgpu_dm_crtc
Alex Hung [Thu, 30 Apr 2026 22:29:06 +0000 (16:29 -0600)]
drm/amd/display: Add KUnit tests for amdgpu_dm_crtc

Add KUnit coverage for functions in amdgpu_dm_crtc.c:
- amdgpu_dm_crtc_modeset_required: verify active+needs_modeset
  combinations (mode_changed, active_changed, connectors_changed)
- amdgpu_dm_crtc_vrr_active_irq: verify all VRR state enum values
- amdgpu_dm_crtc_vrr_active: verify all VRR state enum values
- amdgpu_dm_is_headless: null adev, no connectors, writeback-only,
  disconnected display, connected display, and mixed connector cases
- amdgpu_dm_crtc_helper_mode_fixup: verify it accepts the mode
- amdgpu_dm_crtc_set_vupdate_irq: verify the otg_inst == -1 early
  return using a DRM mock device
- idle_create_workqueue: verify the idle workqueue is allocated and
  initialized in a disabled, non-running state

Assisted-by: Copilot:Claude-Opus-4.8
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Add KUnit tests for amdgpu_dm_irq
Alex Hung [Thu, 30 Apr 2026 21:00:51 +0000 (15:00 -0600)]
drm/amd/display: Add KUnit tests for amdgpu_dm_irq

Add KUnit tests for helper functions, IRQ table management paths, and
DRM mock-backed CRTC lookup in amdgpu_dm_irq.c.

Tests cover:
- amdgpu_dm_hpd_to_dal_irq_source(): all HPD types 1-6,
  AMDGPU_HPD_NONE, and out-of-range values
- are_sinks_equal(): NULL inputs, signal mismatch, EDID
  length mismatch, EDID data mismatch, identical sinks,
  zero-length EDID, full-length identical EDID, and a
  single trailing-byte difference
- dmub_notification_type_str(): notification type mappings that are
  always built, plus the unknown/default case
- amdgpu_dm_irq_init(): low/high handler list initialization
- amdgpu_dm_irq_register_interrupt(): NULL input rejection,
  invalid context/source rejection, low/high handler insertion,
  multiple handlers on one source, and the same handler registered
  in both low and high contexts
- amdgpu_dm_irq_unregister_interrupt(): invalid source and NULL
  handler rejection, removal of registered low/high handlers, and
  the handler-not-found path
- amdgpu_dm_irq_fini(): cleanup of registered low/high handlers and
  the empty-table case
- amdgpu_dm_get_crtc_by_otg_inst(): DRM mock CRTC list match,
  no-match, and empty-list paths

Assisted-by: Copilot:Claude-Opus-4
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Add KUnit tests for amdgpu_dm_connector
Alex Hung [Thu, 30 Apr 2026 04:16:11 +0000 (22:16 -0600)]
drm/amd/display: Add KUnit tests for amdgpu_dm_connector

Add KUnit tests for helper functions in amdgpu_dm_connector.c,
including both pure helper tests and DRM mock-based tests.

Tests cover:
- get_subconnector_type(): all dongle types and unknown default
- get_output_content_type(): all content type mappings and unknown
  default
- adjust_colour_depth_from_display_info(): depth reduction from 12bpc
  to 10bpc, 16bpc no-fallback, YCbCr420 clock halving, and no-fit
  rejection
- get_output_color_space(): RGB full/limited, YCbCr default 709/601,
  BT601/709 with Y_ONLY, OPRGB, BT2020 RGB/YCC paths
- convert_dc_color_depth_into_bpc(): all depths and undefined default
- convert_color_depth_from_display_info(): non-Y420 bpc values, Y420
  default/10/12/16bpc, requested odd bpc rounding, unsupported bpc,
  and requested_bpc capping
- to_drm_connector_type(): HDMI, eDP, LVDS, RGB, DP/MST, DVI single
  and dual link DVII/DVID, virtual, and unknown
- is_duplicate_mode(): empty list, match, no-match, and same-size
  different-clock cases
- amdgpu_dm_get_encoder_crtc_mask(): 1-6 CRTCs and default
- get_aspect_ratio(): all HDMI picture aspect ratios
- decide_crtc_timing_for_drm_display_mode(): scale enabled, matching
  mode, no copy, and no crtc_clock cases
- amdgpu_dm_connector_funcs_reset(): default fields, eDP ABM level
  set, and eDP ABM disabled
- amdgpu_dm_connector_atomic_duplicate_state(): field copy
  verification
- amdgpu_dm_fill_hdr_info_packet(): null metadata early return and
  output zeroing
- amdgpu_dm_connector_atomic_set_property(): scaling center/aspect/
  fullscreen/none/unchanged, underscan hborder/vborder/enable, abm
  sysfs control/level off/level value, and unknown property -EINVAL
- amdgpu_dm_connector_atomic_get_property(): scaling center/aspect/
  full/off, underscan borders, abm sysfs allowed/level/disabled, and
  unknown property -EINVAL
- amdgpu_dm_get_highest_refresh_rate_mode(): null writeback, cached
  base mode, and preferred mode selection
- amdgpu_dm_is_freesync_video_mode(): null mode, match, and no-match
  cases

Assisted-by: Copilot:Claude-Opus-4.8
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Add KUnit tests for amdgpu_dm_dmub
Alex Hung [Thu, 30 Apr 2026 03:28:36 +0000 (21:28 -0600)]
drm/amd/display: Add KUnit tests for amdgpu_dm_dmub

Add KUnit tests for amdgpu_dm_dmub.c covering the following
functions:

- dm_register_dmub_notify_callback(): NULL callback rejection,
  out-of-range type, valid registration with offload flag
- dm_dmub_aux_setconfig_callback(): copy and complete on AUX
  reply, non-AUX skip, NULL dm_notify, SET_CONFIG reply
- dm_dmub_aux_fused_io_callback(): copy reply and complete,
  max ddc_line boundary
- dm_get_default_ips_mode(): IPS mode per DCN version (3.5,
  3.5.1, 3.6, 4.2), disabled for older ASICs, default enabled
  for unhandled newer ASICs
- dm_dmub_hw_init(): early returns for no dmub_srv, no fb_info,
  no firmware
- dm_dmub_hw_resume(): no-op when dmub_srv is NULL
- dm_dmub_sw_init(): returns 0 for unsupported ASIC
- dm_init_microcode(): returns 0 for unsupported ASIC

Assisted-by: Copilot:Claude-Opus-4.6
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Add KUnit tests for amdgpu_dm_audio
Alex Hung [Thu, 30 Apr 2026 03:17:43 +0000 (21:17 -0600)]
drm/amd/display: Add KUnit tests for amdgpu_dm_audio

Add KUnit tests for amdgpu_dm_audio.c.

Tests cover:
- amdgpu_dm_audio_init(): early exit when audio is disabled
- amdgpu_dm_audio_fini(): early exit when audio is not enabled
- fill_audio_info(): manufacturer and product ID propagation,
  display name copy, speaker allocation flags, CEA revision
  gating of audio mode copying (including the zero-mode case),
  and latency field propagation
- amdgpu_dm_audio_component_bind()/unbind(): component ops, device,
  and audio_component pointer are wired up on bind and cleared on
  unbind
- amdgpu_dm_audio_eld_notify(): callback is forwarded with the
  correct port and audio pointer, and the no-op guard paths for a
  missing component, audio_ops, or pin_eld_notify callback

Assisted-by: Copilot:Claude-Opus-4.8
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Add KUnit tests for amdgpu_dm_backlight
Alex Hung [Thu, 30 Apr 2026 02:57:54 +0000 (20:57 -0600)]
drm/amd/display: Add KUnit tests for amdgpu_dm_backlight

Add KUnit tests for the backlight helpers in amdgpu_dm_backlight.c.

Tests cover:
- amdgpu_dm_update_backlight_caps(): short-circuit on populated caps
  and default value assignment
- get_brightness_range(): NULL, PWM-only, and AUX backlight paths
- convert_brightness_to_user(): minimum clamp, maximum passthrough,
  and mid-range rescaling
- convert_brightness_from_user(): linear rescaling, AUX path, and
  custom-curve mapping
- convert_custom_brightness(): exact match, below-first, interpolation,
  above-last, single data point, zero lower luminance, and the
  debug-mask and no-data-point guards
- amdgpu_dm_update_connector_ext_caps(): negative bl_idx and non-eDP
  early returns, OLED defaults, luminance range copy, and the
  amdgpu_backlight force-AUX/force-PWM overrides
- amdgpu_dm_should_create_sysfs(): forced ABM, non-eDP, missing
  backlight index, and AUX vs PWM backlight
- amdgpu_dm_setup_backlight_device(): non-eDP/LVDS skip, disconnected
  link skip, eDP-count limit, and the successful eDP setup path

Assisted-by: Copilot:Claude-Opus-4.8
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Add KUnit tests for amdgpu_dm
Alex Hung [Thu, 30 Apr 2026 21:46:34 +0000 (15:46 -0600)]
drm/amd/display: Add KUnit tests for amdgpu_dm

Add KUnit tests for pure helper functions in amdgpu_dm.c.

Tests cover:
- dm_plane_layer_index_cmp(): equal, ascending, and descending
  layer_index ordering
- fill_plane_color_attributes(): RGB plus BT601/BT709/BT2020
  full- and limited-range YCbCr, and invalid encoding
- modereset_required(): active vs inactive stream states with
  and without a mode change
- dm_get_oriented_plane_size(): 0/90/180/270 degree rotations
- dm_get_plane_scale(): identity, rotated identity, and
  division-by-zero guard
- is_scaling_state_different(): identical state, scaling mode
  change, and underscan enable/border changes
- is_timing_unchanged_for_freesync(): NULL args, identical
  modes, VRR vtotal/vsync shift, and pixel clock change
- set_freesync_fixed_config(): fixed refresh-rate computation
- is_dc_timing_adjust_needed(): pending hw adjust, VRR
  active-fixed, VRR active-state toggle, and steady state
- set_multisync_trigger_params(): disabled trigger and
  rising/falling edge selection by vsync polarity
- set_master_stream(): highest refresh-rate selection and the
  default-to-first-stream case

Assisted-by: Copilot:Claude-Opus-4.8
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Add Support for HDMI Compliance Automation
Fangzhi Zuo [Wed, 3 Jun 2026 17:39:13 +0000 (13:39 -0400)]
drm/amd/display: Add Support for HDMI Compliance Automation

Add support to get DUT trained at FRL link rate when working with
Teledyne M41h compliance automation.

Reviewed-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Fangzhi Zuo <Jerry.Zuo@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Refactor surface_update_flags to flat struct with helpers
Rafal Ostrowski [Wed, 20 May 2026 08:44:17 +0000 (10:44 +0200)]
drm/amd/display: Refactor surface_update_flags to flat struct with helpers

[Why]
The union surface_update_flags type uses a union with a raw
uint32_t member to allow bulk clear/set/test operations on the
bitfield. This couples the struct layout to a specific integer
width, breaks when the number of flag bits exceeds 32, and
scatters raw-access patterns across many call sites. Replacing
the union with a plain struct and adding explicit helper
functions makes the intent clearer and prepares the code for
future flag-set expansion.

[How]
Rename union surface_update_flags to struct pipe_update_bits
and remove the union wrapper, the .bits sub-struct, and the
.raw member. Add inline helpers in dc.h:
surface_update_flags_clear(), surface_update_flags_set_full(),
and surface_update_flags_is_any_set() that operate on the new
struct via memset/memcmp. Add stream_update_flags_clear() and
stream_update_flags_set_full() in dc_stream.h for the stream
update flags union. Update all callers: change the type name,
replace .bits.field with .field, replace .raw = 0 with the
clear helper, replace .raw = 0xFFFFFFFF with the set_full
helper, and replace .raw boolean tests with is_any_set.

Reviewed-by: Nicholas Kazlauskas <nicholas.kazlauskas@amd.com>
Signed-off-by: Rafal Ostrowski <rafal.ostrowski@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Enable pstate for DCN4 non-emulation builds
Gabe Teeger [Tue, 2 Jun 2026 15:38:35 +0000 (11:38 -0400)]
drm/amd/display: Enable pstate for DCN4 non-emulation builds

[Why]
Pstate was disabled during bring-up to avoid interference. Now that
bring-up is complete it can be enabled for non-emulation builds.

[How]
Set pstate_enabled to true in debug_defaults_drv for
non-emulation DCN4 builds.

Reviewed-by: Matthew Stewart <matthew.stewart2@amd.com>
Signed-off-by: Gabe Teeger <gabe.teeger@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Add PSR Active VTotal Control capability
Robin Chen [Sun, 31 May 2026 08:55:26 +0000 (16:55 +0800)]
drm/amd/display: Add PSR Active VTotal Control capability

[WHY]
The PSRSU-RC capability should be populated in DC during edp detection.

Reviewed-by: Aric Cyr <aric.cyr@amd.com>
Signed-off-by: Robin Chen <robin.chen@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Extract connector and encoder code to amdgpu_dm_connector
Alex Hung [Tue, 28 Apr 2026 04:32:01 +0000 (22:32 -0600)]
drm/amd/display: Extract connector and encoder code to amdgpu_dm_connector

Move connector lifecycle functions (init, detect, mode validation,
property handling, EDID parsing, hotplug processing) and encoder
functions (init, destroy, atomic_check, helper_funcs) from amdgpu_dm.c
to amdgpu_dm_connector.c.

No functional change intended.

Assisted-by: Copilot:Claude-Opus-4.6
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Move HPD and IRQ handler code to amdgpu_dm_irq
Alex Hung [Thu, 30 Apr 2026 17:23:59 +0000 (11:23 -0600)]
drm/amd/display: Move HPD and IRQ handler code to amdgpu_dm_irq

Move HPD handling (workqueue creation, debounce, handler registration)
and IRQ handler callbacks (vblank, pflip, vupdate, vline0, outbox) from
amdgpu_dm.c into the existing amdgpu_dm_irq.c. This keeps all
IRQ-related code together rather than creating additional files.

No functional change intended.

Assisted-by: Copilot:Claude-Opus-4.6
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Extract DMUB code to amdgpu_dm_dmub
Alex Hung [Tue, 28 Apr 2026 02:49:30 +0000 (20:49 -0600)]
drm/amd/display: Extract DMUB code to amdgpu_dm_dmub

Move DMUB-related functions and firmware defines from amdgpu_dm.c
into new amdgpu_dm_dmub.c and amdgpu_dm_dmub.h files to reduce
the size of amdgpu_dm.c and improve code organization.

No functional change intended.

Assisted-by: Copilot:Claude-Opus-4.6
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Extract audio code to amdgpu_dm_audio
Alex Hung [Tue, 28 Apr 2026 01:20:53 +0000 (19:20 -0600)]
drm/amd/display: Extract audio code to amdgpu_dm_audio

Move audio component, init/fini, ELD notification,
fill_audio_info, and commit_audio functions from
amdgpu_dm.c into a dedicated amdgpu_dm_audio.c file
with its own header.

No functional change intended.

Assisted-by: Copilot:Claude-Opus-4.6
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Extract backlight code to amdgpu_dm_backlight
Alex Hung [Fri, 24 Apr 2026 00:34:52 +0000 (18:34 -0600)]
drm/amd/display: Extract backlight code to amdgpu_dm_backlight

Move backlight-related functions from amdgpu_dm.c into a new
amdgpu_dm_backlight.c file to improve code organization and
reduce the size of the monolithic amdgpu_dm.c.

No functional change intended.

Assisted-by: Copilot:Claude-Opus-4.6
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Introduce dc_plane_cm and migrate surface update color path
Rafal Ostrowski [Fri, 22 May 2026 06:02:16 +0000 (08:02 +0200)]
drm/amd/display: Introduce dc_plane_cm and migrate surface update color path

[Why]
Begin convergence with upstream Color Manager refactor
(fda768acb2a1 "drm/amd/display: Sync dcn42 with DC 3.2.373") by
consolidating fragmented per-plane CM state (shaper, 3DLUT, blend,
CM2) into a single dc_plane_cm structure shared by dc_plane_state
and dc_surface_update. Legacy fields are gated behind TRIM_CM2 so
that it keeps compatibility with other repositories.

[How]
Refactored to use newer structures.
No functional behavior change intended. Under !TRIM_CM2 the legacy
fields are still populated for compatibility with other repositories.

v2: squash in conflicting types fix

Reviewed-by: Dillon Varone <dillon.varone@amd.com>
Signed-off-by: Rafal Ostrowski <rafal.ostrowski@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Remove unused project_id from DML2 core instance
Wenjing Liu [Mon, 1 Jun 2026 21:40:18 +0000 (17:40 -0400)]
drm/amd/display: Remove unused project_id from DML2 core instance

[Why]
The project_id field stored in dml2_core_instance and related
context structs was not consumed after initial setup and
represents unnecessary coupling between the core layer and
project-specific identifiers.

[How]
- Remove project_id field from dml2_core_instance
- Remove the corresponding assignment in dml2_core_create

Reviewed-by: Austin Zheng <austin.zheng@amd.com>
Signed-off-by: Wenjing Liu <wenjing.liu@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Remove get_utm_qos_model from soc_and_ip_translator
Wenjing Liu [Thu, 28 May 2026 20:43:26 +0000 (16:43 -0400)]
drm/amd/display: Remove get_utm_qos_model from soc_and_ip_translator

[Why]
The QoS model is now populated directly in clock manager
from firmware data. The translator function pointer is no
longer needed.

[How]
- Remove get_utm_qos_model function pointer from
  soc_and_ip_translator_funcs
- Remove associated forward declarations from
  soc_and_ip_translator.h

Reviewed-by: Dillon Varone <dillon.varone@amd.com>
Signed-off-by: Wenjing Liu <wenjing.liu@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Add utm_qos_model pointer to clk_bw_params
Wenjing Liu [Thu, 28 May 2026 17:10:27 +0000 (13:10 -0400)]
drm/amd/display: Add utm_qos_model pointer to clk_bw_params

[Why]
Add support for passing QoS model data from clock manager
to bandwidth calculation consumers.

[How]
- Add forward declaration and const pointer for utm_qos_model
  in clk_bw_params

Reviewed-by: Dillon Varone <dillon.varone@amd.com>
Signed-off-by: Wenjing Liu <wenjing.liu@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Add a new interface to set idle opts in clock manager
Nicholas Kazlauskas [Thu, 28 May 2026 14:51:20 +0000 (10:51 -0400)]
drm/amd/display: Add a new interface to set idle opts in clock manager

[Why & How]
For future use in migrating the idle optimizations message to PMFW to
DC core.

Reviewed-by: Dillon Varone <dillon.varone@amd.com>
Signed-off-by: Nicholas Kazlauskas <nicholas.kazlauskas@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Increase dcn42b uclk value
Gabe Teeger [Thu, 28 May 2026 21:33:16 +0000 (17:33 -0400)]
drm/amd/display: Increase dcn42b uclk value

Increase uclk value in order to enable UHBR20.

Reviewed-by: Dillon Varone <dillon.varone@amd.com>
Signed-off-by: Gabe Teeger <gabe.teeger@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu/ras: Add address sanity check for uniras
Ce Sun [Thu, 11 Jun 2026 07:38:46 +0000 (15:38 +0800)]
drm/amdgpu/ras: Add address sanity check for uniras

Add address sanity check for uniras

Signed-off-by: Ce Sun <cesun102@amd.com>
Reviewed-by: Tao Zhou <tao.zhou1@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu: correct reservation fence slots for userq per-vm BOs eviction
Prike Liang [Thu, 11 Jun 2026 02:58:05 +0000 (10:58 +0800)]
drm/amdgpu: correct reservation fence slots for userq per-vm BOs eviction

It fixes both the move overflow and the eviction fence add for
evicting these per-vm BOs.

Signed-off-by: Prike Liang <Prike.Liang@amd.com>
Reviewed-by: Christian König <christian.koenig@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdkfd: add sdma queue counter for gfxv9.4.3
Eric Huang [Thu, 4 Jun 2026 13:24:32 +0000 (09:24 -0400)]
drm/amdkfd: add sdma queue counter for gfxv9.4.3

since gfx 9.4.3 HW is calculating accumulated activity counter
per-queue in register sdmax_rlcx_utilization_hi/lo, CPFW adds it in
sdma MQD for save/restore, KFD will read it from there. gfx 9.4.2
will still keep the way to read from memory at rptr+8.

v2: read dynamic counter directly from utilization register
v3: add CPFW supported version check (Harish)

Signed-off-by: Eric Huang <jinhuieric.huang@amd.com>
Reviewed-by: Harish Kasiviswanathan <Harish.Kasiviswanathan@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdkfd: Add gfx11 queue/pipe reset support to topology
Amber Lin [Fri, 5 Jun 2026 22:18:10 +0000 (18:18 -0400)]
drm/amdkfd: Add gfx11 queue/pipe reset support to topology

Add gfx11 queue/pipe reset support to KFD topology

Signed-off-by: Amber Lin <Amber.Lin@amd.com>
Reviewed-by: Jesse Zhang <jesse.zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdkfd: Fix reset event signal
Amber Lin [Tue, 9 Jun 2026 16:33:40 +0000 (12:33 -0400)]
drm/amdkfd: Fix reset event signal

During the KFD/KCQ coordination rework, bad queues not requiring reset
were combined into the rework and generated wrong reset signals to the
process. Fix it by adding the reset check.

Signed-off-by: Amber Lin <Amber.Lin@amd.com>
Reviewed-by: Shaoyun Liu <shaoyun.liu@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu: Don't use UTS_RELEASE directly
Uwe Kleine-König (The Capable Hub) [Tue, 28 Apr 2026 14:47:03 +0000 (16:47 +0200)]
drm/amdgpu: Don't use UTS_RELEASE directly

UTS_RELEASE evaluates to a static string and changes quite easily (e.g.
uncommitted changes in the source tree or new commits). So when checking
if a patch introduces changes to the resulting binary each usage of
UTS_RELEASE is source of annoyance.

Instead of using UTS_RELEASE directly use init_utsname()->release which
evaluates to the same string but with that a change of UTS_RELEASE
doesn't affect amdgpu_dev_coredump.o.

Reviewed-by: Mario Limonciello (AMD) <superm1@kernel.org>
Signed-off-by: Uwe Kleine-König (The Capable Hub) <u.kleine-koenig@baylibre.com>
Link: https://patch.msgid.link/20260428144704.1114562-2-u.kleine-koenig@baylibre.com
Signed-off-by: Mario Limonciello <mario.limonciello@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu/gfx: fix cleaner shader IB buffer overflow
Asad Kamal [Fri, 5 Jun 2026 15:44:08 +0000 (23:44 +0800)]
drm/amdgpu/gfx: fix cleaner shader IB buffer overflow

The cleaner shader sysfs path allocates a 16-dword (64 byte) IB but
incorrectly fills (align_mask + 1) dwords. On GFX rings align_mask is
0xff, so the loop wrote 256 dwords into a 64-byte buffer, causing a
kernel page fault.

The IB only needs to be a minimal NOP shell to schedule the job; the
cleaner shader itself is emitted on the ring via emit_cleaner_shader().
Fill 16 dwords to match the allocation.

v2: Use ib_size_dw variable (Lijo)

Fixes: d361ad5d2fc0 ("drm/amdgpu: Add sysfs interface for running cleaner shader")
Suggested-by: Lijo Lazar <lijo.lazar@amd.com>
Signed-off-by: Asad Kamal <asad.kamal@amd.com>
Reviewed-by: Lijo Lazar <lijo.lazar@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu/ras: Add flag to make VBIOS read optional
Ce Sun [Sat, 6 Jun 2026 13:20:23 +0000 (21:20 +0800)]
drm/amdgpu/ras: Add flag to make VBIOS read optional

Add flag to make VBIOS read optional

Signed-off-by: Ce Sun <cesun102@amd.com>
Reviewed-by: Hawking Zhang <Hawking.Zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu: allocate lockdep mutex on the heap to fix stack overflow
Prike Liang [Fri, 5 Jun 2026 07:28:40 +0000 (15:28 +0800)]
drm/amdgpu: allocate lockdep mutex on the heap to fix stack overflow

Replace the stack-allocated amdgpu_lockdep mutex with a heap allocation
via kmalloc to fix a stack overflow caused by the large struct size.

Signed-off-by: Prike Liang <Prike.Liang@amd.com>
Reviewed-by: Vitaly Prosyak <vitaly.prosyak@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu/ras: added RAS EEPROM device support check
Ce Sun [Sat, 6 Jun 2026 13:15:54 +0000 (21:15 +0800)]
drm/amdgpu/ras: added RAS EEPROM device support check

Added RAS EEPROM device support check

Signed-off-by: Ce Sun <cesun102@amd.com>
Reviewed-by: Hawking Zhang <Hawking.Zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu/ras: Parse all deferred errors with UMC aca handle
Ce Sun [Wed, 3 Jun 2026 02:45:48 +0000 (10:45 +0800)]
drm/amdgpu/ras: Parse all deferred errors with UMC aca handle

We should only increase the deferred errors in UMC block

Signed-off-by: Ce Sun <cesun102@amd.com>
Reviewed-by: Tao Zhou <tao.zhou1@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Promote DC to 3.2.385
Taimur Hassan [Sat, 30 May 2026 00:40:42 +0000 (19:40 -0500)]
drm/amd/display: Promote DC to 3.2.385

Summary:

  * Display connectivity & HPD:
    - Retry link detection on resume, boot, and hotplug
    - Refactor HPD RX to use handle_hpd_irq_helper with detect reason
    - Always create delayed HPD work queue
    - Restore periodic detection for DCN35

  * DCN42B support:
    - Fix DCN42B version detection
    - Add DCN42B to dml21_translation_helper

  * KUnit testing infrastructure:
    - Add KUnit tests for amdgpu_dm_pp_smu, amdgpu_dm_mst_types,
      and writeback connector
    - Extract HDCP and DPRX CRC transition helpers for KUnit
    - Export symbols for KUnit test modules
    - Enable warnings as errors for KUnit tests

  * Fixes & cleanups:
    - Fix compressed buffer config routine waiting time
    - Fix incorrect logic in CRC source handling
    - Fix writeback format loop and variable init
    - Fix max dispclk_khz/dppclk_khz double 1000
    - Remove duplicate pp_rn_set_wm_ranges
    - Remove dead code in dm_dp_mst_get_modes
    - Remove redundant code in amdgpu_dm_replay
    - Skip PHY SSC reduction on some 8K panels
    - Temp disable repeater FGCG as workaround
    - Deprecate DMUB register offload functionality
    - TEST_HARNESS FSN could be 0

  * Firmware:
    - DMUB FW promotion to 0.1.62.0

Signed-off-by: Taimur Hassan <Syed.Hassan@amd.com>
Signed-off-by: Aurabindo Pillai <aurabindo.pillai@amd.com>
Reviewed-by: Alex Hung <alex.hung@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: fix compressed buffer config routine waiting time
Antonio Quartulli [Tue, 19 May 2026 15:57:28 +0000 (15:57 +0000)]
drm/amd/display: fix compressed buffer config routine waiting time

Replace the four open-coded REG_WAIT calls with calls to
dcn31_wait_for_det_apply() so the compressed buffer (compbuf) sizing
path waits long enough for the DET size update to take effect, and the
wait timing stays consistent across the driver.

No functional change beyond the corrected timeout.

Signed-off-by: Antonio Quartulli <antonio@mandelbit.com>
Signed-off-by: Aurabindo Pillai <aurabindo.pillai@amd.com>
Reviewed-by: Alex Hung <alex.hung@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Use handle_hpd_irq_helper for HPD RX
Timur Kristóf [Sun, 31 May 2026 10:57:41 +0000 (12:57 +0200)]
drm/amd/display: Use handle_hpd_irq_helper for HPD RX

Remove duplicated code and just call handle_hpd_irq_helper
with the appropriate detect reason.

Signed-off-by: Timur Kristóf <timur.kristof@gmail.com>
Signed-off-by: Aurabindo Pillai <aurabindo.pillai@amd.com>
Reviewed-by: Alex Hung <alex.hung@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdkfd: Fix SMI event PID reporting for containers
Andrew Martin [Thu, 28 May 2026 14:32:52 +0000 (10:32 -0400)]
drm/amdkfd: Fix SMI event PID reporting for containers

SMI events were reporting incorrect PIDs in containerized environments,
causing test failures where container processes expected to see their
namespace-local PIDs but instead received global host PIDs.

The issue had two root causes:

1. Event functions were called from kernel context (page fault handlers,
   migration workers) where 'current' refers to the kernel worker thread,
   not the userspace GPU process that triggered the event.

2. PID conversion used task_tgid_vnr() which returns the PID in the
   caller's namespace (init namespace for kernel threads), not the task's
   own namespace.

This patch updates the SMI event interface:

- Change 8 event function signatures to accept task_struct pointer
  instead of pid_t, allowing proper namespace-aware PID conversion

- Convert PIDs using task_tgid_nr_ns(task, task_active_pid_ns(task))
  which returns the PID as the process sees it via getpid()

- Update 10 call sites to pass p->lead_thread (the GPU process)
  instead of p->lead_thread->pid or current (kernel worker)

This ensures SMI events report container-local PIDs, which is critical
for containerized GPU workloads to correctly correlate events with their
processes.

Tested-by: Andrew Martin <andmarti@amd.com>
Assisted-by: Claude:Sonnet 4-5
Signed-off-by: Andrew Martin <andrew.martin@amd.com>
Reviewed-by: Harish Kasiviswanathan <Harish.Kasiviswanathan@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Add detect reason to handle_hpd_irq_helper
Timur Kristóf [Sun, 31 May 2026 10:57:40 +0000 (12:57 +0200)]
drm/amd/display: Add detect reason to handle_hpd_irq_helper

This makes it possible to reuse the function for other purposes
in the next few commits, such as HPD RX.

Signed-off-by: Timur Kristóf <timur.kristof@gmail.com>
Signed-off-by: Aurabindo Pillai <aurabindo.pillai@amd.com>
Reviewed-by: Alex Hung <alex.hung@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu/gfx: defer per-queue helper_end until after MES resume
Jesse Zhang [Fri, 5 Jun 2026 08:28:47 +0000 (16:28 +0800)]
drm/amdgpu/gfx: defer per-queue helper_end until after MES resume

amdgpu_gfx_reset_mes_compute() runs amdgpu_mes_suspend(adev, 0) to
quiesce all gangs, resets the offending queue(s), then resumes. The
existing amdgpu_gfx_mes_reset_queue() called amdgpu_ring_reset_helper_end()
right after unmap/restore/map of the reset queue, which re-emits backed-up
commands and rings the doorbell. That doorbell hits a still-suspended CP:
on the subsequent resume the queue partially wedges -- the first new IB
after the reset may execute but later submissions stall, which surfaces
as repeated timeouts on the same ring under concurrent workloads.

Split out amdgpu_gfx_mes_reset_queue_start() (backup + MES reset +
unmap/restore/map only) and defer helper_end. amdgpu_gfx_reset_mes_compute()
collects the (ring, fence) pair for every queue it resets and runs
helper_end on each after amdgpu_mes_resume(), so the re-emit doorbells
land on a running CP. amdgpu_gfx_reset_mes_kcq() now reports the matched
ring/fence back to the caller for the same reason.

Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Jesse Zhang <jesse.zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Add KUnit tests for amdgpu_dm_pp_smu
Alex Hung [Fri, 29 May 2026 16:06:22 +0000 (10:06 -0600)]
drm/amd/display: Add KUnit tests for amdgpu_dm_pp_smu

[WHAT]
Add KUnit tests for two functions in amdgpu_dm_pp_smu.c:
get_default_clock_levels and dc_to_pp_clock_type.

Assisted-by: Copilot:Claude-Opus-4.6
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Aurabindo Pillai <aurabindo.pillai@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Remove duplicate pp_rn_set_wm_ranges
Alex Hung [Mon, 25 May 2026 21:27:00 +0000 (15:27 -0600)]
drm/amd/display: Remove duplicate pp_rn_set_wm_ranges

[WHAT]
Remove pp_rn_set_wm_ranges and reuse the identical
pp_nv_set_wm_ranges for the DCN_VERSION_2_1 case instead.

Assisted-by: Copilot:Claude-Opus-4.6
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Aurabindo Pillai <aurabindo.pillai@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Restore periodic detection for DCN35
Ivan Lipski [Thu, 28 May 2026 16:28:51 +0000 (12:28 -0400)]
drm/amd/display: Restore periodic detection for DCN35

[Why&How]
Periodic detection callbacks from DCN35 was removed for higher IPS
residency causing some displays to fail to recover after DPMS sleep. The
monitors bounces HPD ~1.2s after link training, and without periodic
detection the system enters IPS with no mechanism to wake and rediscover
the display.

Restore the periodic detection calls in dcn35_clk_mgr for now. It should
be replaced with a proper IPS-aware solution long term using DMUB.

Also remove it from dcn31 and dcn314_clk_mgr.c since they do not have IPS,
thus should not affect them.

Fixes: 3f6c060846be ("drm/amd/display: Remove periodic detection callbacks from dcn35+")
Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5318
Reviewed-by: Nicholas Kazlauskas <nicholas.kazlauskas@amd.com>
Signed-off-by: Ivan Lipski <ivan.lipski@amd.com>
Signed-off-by: Aurabindo Pillai <aurabindo.pillai@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Extract HDCP testable helpers for KUnit
Alex Hung [Wed, 27 May 2026 23:01:13 +0000 (17:01 -0600)]
drm/amd/display: Extract HDCP testable helpers for KUnit

[WHAT]
Extract hdcp_get_content_protection_from_status() and
hdcp_get_link_display_adjustments() from event_property_update()
and hdcp_update_display() so the pure decision logic can be
KUnit-tested.

Also update function comments to kernel-doc formats.

Assisted-by: Copilot:Claude-Opus-4.6
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Aurabindo Pillai <aurabindo.pillai@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Extract DPRX CRC transition helpers for KUnit testing
Alex Hung [Thu, 28 May 2026 20:01:08 +0000 (14:01 -0600)]
drm/amd/display: Extract DPRX CRC transition helpers for KUnit testing

Extract three pure predicate functions from amdgpu_dm_crtc_set_crc_source():
- dm_need_dp_aux
- dm_crc_source_should_start_dprx
- dm_crc_source_should_stop_dprx

Refactor set_crc_source() to use these helpers, replacing the nested
if/else if structure with flat, mutually-exclusive branches driven by
the new predicates.

Add KUnit test cases covering all relevant source combinations for each
helper, including the regression case where DPRX→NONE must trigger
drm_dp_stop_crc().

Assisted-by: Copilot:Claude-Sonnet-4.6
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Aurabindo Pillai <aurabindo.pillai@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Fix incorrect logic in CRC source handling
Alex Hung [Thu, 28 May 2026 17:48:11 +0000 (11:48 -0600)]
drm/amd/display: Fix incorrect logic in CRC source handling

[WHAT]
Fix three issues amdgpu_dm_crc.c:
- Use cur_crc_src instead of source when deciding whether to call
  drm_dp_stop_crc() in the disable path of set_crc_source(). When
  disabling CRC, source is always NONE so dm_is_crc_source_dprx(source)
  was always false, meaning drm_dp_stop_crc() was never called when
  stopping a DPRX CRC source. Use cur_crc_src to check what was
  previously active instead.
- Replace fragile 'source < 0' comparisons in verify_crc_source() and
  set_crc_source() with AMDGPU_DM_PIPE_CRC_SOURCE_INVALID.
  and avoiding signed/unsigned enum comparison concerns.
- Remove redundant NULL initializations for drm_dev and acrtc in
  handle_crc_irq(). Both variables are unconditionally assigned right
  after.

Assisted-by: Copilot:Claude-Sonnet-4.6
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Aurabindo Pillai <aurabindo.pillai@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Add KUnit tests for amdgpu_dm_mst_types
Alex Hung [Mon, 25 May 2026 20:48:34 +0000 (14:48 -0600)]
drm/amd/display: Add KUnit tests for amdgpu_dm_mst_types

[WHAT]
Add KUnit test coverage for needs_dsc_aux_workaround() in
amdgpu_dm_mst_types.c. Tests verify the function correctly
identifies links requiring the DSC AUX workaround based on
branch device ID, DPCD revision, and sink count.

Assisted-by: Copilot:Claude-Opus-4.6
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Aurabindo Pillai <aurabindo.pillai@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Remove dead code in dm_dp_mst_get_modes
Alex Hung [Mon, 25 May 2026 19:12:16 +0000 (13:12 -0600)]
drm/amd/display: Remove dead code in dm_dp_mst_get_modes

[WHAT]
Remove unreachable null check on aconnector after container_of,
and redundant dc_sink checks where dc_sink is guaranteed non-NULL
after earlier null-check with early return.

Assisted-by: Copilot:Claude-Opus-4.6
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Aurabindo Pillai <aurabindo.pillai@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Enable warnings as errors for KUnit tests
Alex Hung [Wed, 27 May 2026 23:35:47 +0000 (17:35 -0600)]
drm/amd/display: Enable warnings as errors for KUnit tests

[WHAT]
Add CONFIG_WERROR=y to .kunitconfig to treat compiler warnings
as errors during KUnit builds, ensuring warnings are caught
early.

Assisted-by: Copilot:Claude-Opus-4.6
Reviewed-by: Ray Wu <ray.wu@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Aurabindo Pillai <aurabindo.pillai@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: remove redundant code in amdgpu_dm_replay
Alex Hung [Wed, 27 May 2026 22:20:29 +0000 (16:20 -0600)]
drm/amd/display: remove redundant code in amdgpu_dm_replay

[WHAT]
In amdgpu_dm_link_setup_replay(), nom_coasting_vtotal was
used only once immediately after in set_replay_coasting_vtotal().
Inline the value directly to remove the no-op alias.

In amdgpu_dm_set_replay_caps(), replace link->ctx->dc->debug
with dc->debug since dc is already assigned as link->ctx->dc,
eliminating a redundant pointer round-trip.

Assisted-by: Copilot:Claude-Sonnet-4.6
Reviewed-by: Ray Wu <ray.wu@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Aurabindo Pillai <aurabindo.pillai@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: fix max dispclk_khz/dppclk_khz double 1000
Charlene Liu [Fri, 22 May 2026 00:36:01 +0000 (20:36 -0400)]
drm/amd/display: fix max dispclk_khz/dppclk_khz double 1000

[why]
Fix regresson caused by double roundup and index out of range

Reviewed-by: Dillon Varone <dillon.varone@amd.com>
Reviewed-by: Dmytro Laktyushkin <dmytro.laktyushkin@amd.com>
Signed-off-by: Charlene Liu <Charlene.Liu@amd.com>
Signed-off-by: Aurabindo Pillai <aurabindo.pillai@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Add KUnit tests for writeback connector
Alex Hung [Mon, 25 May 2026 18:08:49 +0000 (12:08 -0600)]
drm/amd/display: Add KUnit tests for writeback connector

[WHAT]
Add KUnit tests for amdgpu_dm_wb_encoder_atomic_check() and
amdgpu_dm_wb_connector_get_modes(). Tests cover null job,
null fb, size mismatch, format validation, and mode count
bounds using DRM KUnit mock devices.

Assisted-by: Copilot:Claude-Opus-4.6
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Aurabindo Pillai <aurabindo.pillai@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Fix writeback format loop and variable init
Alex Hung [Mon, 25 May 2026 18:08:20 +0000 (12:08 -0600)]
drm/amd/display: Fix writeback format loop and variable init

[WHAT]
1. Use ARRAY_SIZE() instead of manual sizeof division for the
   format array iteration. Add a break statement to exit the loop
   early once a matching format is found.
2. Remove redundant zero initialization of res since all paths
   assign before use.

Assisted-by: Copilot:Claude-Opus-4.6
Reviewed-by: Bhawanpreet Lakha <bhawanpreet.lakha@amd.com>
Signed-off-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Aurabindo Pillai <aurabindo.pillai@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Temp disable repeater FGCG as workaround
Ovidiu Bunea [Thu, 21 May 2026 19:27:11 +0000 (15:27 -0400)]
drm/amd/display: Temp disable repeater FGCG as workaround

[why & how]
There is an issue that is seemingly limited to DCN42 where systems with
IOMMU enabled will hang during reboot stress testing. The hang happens shortly
after DCN PG exit happens and HUBP is programmed for the first flip, but before
the first surface address is latched. Testing has shown that disabling
DCCG_GLOBAL_FGCG_REP_DIS, HUBP_FGCG_REP_DIS, and DCFCLK_GATE_DIS can mask this
issue.

Disable FGCG for these three repeater bits to avoid issue while debug is on-going.

Reviewed-by: Nicholas Kazlauskas <nicholas.kazlauskas@amd.com>
Signed-off-by: Ovidiu Bunea <ovidiu.bunea@amd.com>
Signed-off-by: Aurabindo Pillai <aurabindo.pillai@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Deprecate DMUB register offload functionality
Austin Zheng [Tue, 12 May 2026 20:20:54 +0000 (16:20 -0400)]
drm/amd/display: Deprecate DMUB register offload functionality

[Why]
The DMUB register offload feature should no longer be used.
This was originally a debug feature for DCN21.
No longer applicable to the DMUB programming model.

[How]
Remove DMUB register offload infrastructure including helper
functions, structures, debug options, and register sequence macros.

Reviewed-by: Nicholas Kazlauskas <nicholas.kazlauskas@amd.com>
Signed-off-by: Austin Zheng <Austin.Zheng@amd.com>
Signed-off-by: Aurabindo Pillai <aurabindo.pillai@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: TEST_HARNESS FSN could be 0
ChunTao Tso [Mon, 23 Mar 2026 05:53:26 +0000 (13:53 +0800)]
drm/amd/display: TEST_HARNESS FSN could be 0

The frame skipping number could be 0 if needed.

Reviewed-by: Robin Chen <robin.chen@amd.com>
Signed-off-by: ChunTao Tso <ChunTao.Tso@amd.com>
Signed-off-by: Aurabindo Pillai <aurabindo.pillai@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/display: Skip PHY SSC reduction on some 8K panels
Roman Li [Wed, 20 May 2026 20:50:34 +0000 (16:50 -0400)]
drm/amd/display: Skip PHY SSC reduction on some 8K panels

[Why]
Some 8K displays cannot tolerate the reduced phy ssc value
at high link utilization and show corruption or black screen.

[How]
Add an EDID panel-id quirk to utilize existing skip_phy_ssc_reduction flag.

To pass the link into the quirk handler, change the signature of
apply_edid_quirks() to take link as an argument. The dev local in
dm_helpers_parse_edid_caps() becomes unused and is removed.

Fixes: 5fa62c87cffd ("drm/amd/display: Add option to disable PHY SSC reduction on transmitter enable")
Reviewed-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Roman Li <Roman.Li@amd.com>
Signed-off-by: Aurabindo Pillai <aurabindo.pillai@amd.com>
Tested-by: Dan Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdkfd: Extend MQDs in HBM to gfx944
Kent Russell [Fri, 8 May 2026 21:12:26 +0000 (17:12 -0400)]
drm/amdkfd: Extend MQDs in HBM to gfx944

This has proven stable and performant on gfx943 and gfx950, so extend
it to gfx944 as well

Signed-off-by: Kent Russell <kent.russell@amd.com>
Reviewed-by: David Francis <David.Francis@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdkfd: Extend MQDs in HBM to gfx942
Kent Russell [Fri, 8 May 2026 21:12:17 +0000 (17:12 -0400)]
drm/amdkfd: Extend MQDs in HBM to gfx942

This has proven stable and performant on gfx943 and gfx950, so extend
it to the Aldebaran/gfx942 series

Signed-off-by: Kent Russell <kent.russell@amd.com>
Reviewed-by: David Francis <David.Francis@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu: remove spurious line in amdgpu_ring_find_guilty_fence()
Alex Deucher [Thu, 4 Jun 2026 20:51:48 +0000 (16:51 -0400)]
drm/amdgpu: remove spurious line in amdgpu_ring_find_guilty_fence()

Copy-paste error.

Fixes: 36ed61b1c01a ("drm/amdgpu/fence: add helper to extract the guilty fence")
Reviewed-by: Jesse Zhang <jesse.zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu: skip already suspended IP blocks in ip_suspend_phase2
Yunxiang Li [Fri, 5 Jun 2026 12:59:34 +0000 (08:59 -0400)]
drm/amdgpu: skip already suspended IP blocks in ip_suspend_phase2

The GPU reload test (S3 / mode1 reset / module reload) triggers a
WARN_ON in amdgpu_irq_put() on gfx10 when unloading amdgpu:

  WARNING: CPU: 0 PID: 2314 at amd/amdgpu/amdgpu_irq.c:676 amdgpu_irq_put+0xc3/0xe0 [amdgpu]
  Call Trace:
   gfx_v10_0_hw_fini+0x41/0x150 [amdgpu]
   amdgpu_ip_block_hw_fini+0x29/0xc0 [amdgpu]
   amdgpu_device_fini_hw+0x315/0x610 [amdgpu]
   amdgpu_driver_unload_kms+0x7c/0x90 [amdgpu]
   amdgpu_pci_remove+0x51/0x90 [amdgpu]

amdgpu_device_ip_resume_phase2() skips IP blocks whose status.hw is
already set, but amdgpu_device_ip_suspend_phase2() never had the
matching guard, so a block can be suspended twice (e.g. a reset or
recovery issued while the device is already suspended).  The second
suspend runs hw_fini again, which now releases the gfx fault IRQs
unconditionally, dropping a refcount that is already zero and tripping
the WARN_ON in amdgpu_irq_put().

The fault/EOP IRQ get/put were balanced through late_init/hw_fini
before, which masked the double-suspend; moving the get into hw_init
made the suspend/resume asymmetry visible as an IRQ refcount underflow.

Honor status.hw in ip_suspend_phase2() so suspend mirrors resume and a
block is only torn down once.

Fixes: 9117d8be850b ("drm/amdgpu/gfx: move fault and EOP IRQ get/put to hw_init/hw_fini")
Fixes: 482f0e538580 ("drm/amdgpu: fix double ucode load by PSP(v3)")
Signed-off-by: Yunxiang Li <Yunxiang.Li@amd.com>
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdkfd: Move mqd_on_vram out of v9 mqd manager
Kent Russell [Mon, 20 Apr 2026 15:19:16 +0000 (11:19 -0400)]
drm/amdkfd: Move mqd_on_vram out of v9 mqd manager

This will allow it to be used outside of gfx9

Signed-off-by: Kent Russell <kent.russell@amd.com>
Reviewed-by: David Francis <David.Francis@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdkfd: Properly acquire queue buffers in CRIU restore
David Francis [Thu, 4 Jun 2026 19:04:03 +0000 (15:04 -0400)]
drm/amdkfd: Properly acquire queue buffers in CRIU restore

When kfd_queue_acquire_buffers() was split off from
set_queue_properties_from_user(), set_queue_properties_from_criu()
was missed. Thus, set_queue_properties_from_criu() is not
filling out the buffer fields of queue_properties, which
can come up when subsequent code expects them to be non-null.

Add the proper call to kfd_queue_acquire_buffers(), and also
use the right cast types in set_queue_properties_from_criu()
(which were missed at the same time)

Signed-off-by: David Francis <David.Francis@amd.com>
Reviewed-by: Kent Russell <kent.russell@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/pm: re-enable MC access after PrepareMp1ForUnload on SMU V15 APUs
Shubhankar Milind Sardeshpande [Thu, 21 May 2026 05:25:18 +0000 (10:55 +0530)]
drm/amd/pm: re-enable MC access after PrepareMp1ForUnload on SMU V15 APUs

During smu_v15_0_0_system_features_control(), the driver sends a
PrepareMp1ForUnload message to PMFW. PMFW then performs nBIF and SYSHUB
function-level resets (FLR), disabling PCIe CFG space reset, which
clears the framebuffer enable bit to zero and disables MC (memory controller)
access from the host.

Re-enable MC access via the nbio mc_access_enable callback right after
PrepareMp1ForUnload completes in smu_v15_0_0_system_features_control().

Signed-off-by: Shubhankar Milind Sardeshpande <Shubhankar.MilindSardeshpande@amd.com>
Signed-off-by: Suresh Guttula <Suresh.Guttula@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu/vcn4.0.5: enable secure submission on unified ring
Jeevana Muthyala [Mon, 25 May 2026 06:19:24 +0000 (11:49 +0530)]
drm/amdgpu/vcn4.0.5: enable secure submission on unified ring

Set secure_submission_supported = true for the VCN unified ring funcs in
vcn_v4_0_5.c so secure IBs are allowed on the unifiedring.
Without this, protected decode submissions are blocked by the
common IB gate and can fail playback for secure content.

For vcn_v4_0_5.c (fixed STX VCN version), secure submission is
enabled directly in the ring funcs definition.

This change only advertises existing hardware/firmware capability;
non-secure decode paths are unaffected.

Signed-off-by: Jeevana Muthyala <jmuthyal@amd.com>
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu/vcn4.0: enable secure submission on unified ring
Jeevana Muthyala [Mon, 25 May 2026 06:13:40 +0000 (11:43 +0530)]
drm/amdgpu/vcn4.0: enable secure submission on unified ring

Set secure_submission_supported = true for the VCN unified ring funcs in
vcn_v4_0.c so secure IBs are allowed on the unified ring.
Without this, protected decode submissions are blocked by the
common IB gate and can fail playback for secure content.

For vcn_v4_0.c, the secure ring funcs are selected for the secure-capable
IP version.

This change only advertises existing hardware/firmware capability;
non-secure decode paths are unaffected.

Signed-off-by: Jeevana Muthyala <jmuthyal@amd.com>
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/pm: sleep on PMFW EEPROM busy in bad page count query
Candice Li [Wed, 3 Jun 2026 01:53:01 +0000 (09:53 +0800)]
drm/amd/pm: sleep on PMFW EEPROM busy in bad page count query

Use usleep_range() instead of mdelay() to match the behavior
of ras_fw_get_badpage_count() in rascore path.

Signed-off-by: Candice Li <candice.li@amd.com>
Reviewed-by: Yang Wang <kevinyang.wang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/ras: sleep on PMFW EEPROM busy in bad page count query
Candice Li [Fri, 29 May 2026 04:29:52 +0000 (12:29 +0800)]
drm/amd/ras: sleep on PMFW EEPROM busy in bad page count query

Use usleep_range() instead of mdelay() when ras_fw_get_badpage_count()
retries on -EBUSY so the driver yields the CPU while waiting for PMFW
EEPROM to become ready.

Signed-off-by: Candice Li <candice.li@amd.com>
Reviewed-by: Yang Wang <kevinyang.wang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu: initialize iter.start in amdgpu_devcoredump_format
Qiang Yu [Tue, 26 May 2026 06:45:48 +0000 (14:45 +0800)]
drm/amdgpu: initialize iter.start in amdgpu_devcoredump_format

This fixes read /sys/class/drm/cardN/device/devcoredump/data
return empty content sometimes.

amdgpu_devcoredump_format() leaves struct drm_print_iterator's
.start field uninitialized on the stack before passing it to
drm_coredump_printer(). __drm_puts_coredump() compares the running
.offset against .start to decide whether to skip or copy each
chunk:

if (iterator->offset < iterator->start) {
if (iterator->offset + len <= iterator->start) {
iterator->offset += len;
return;
}
...
}

Fixes: 4bbba79a7f1d ("drm/amdgpu: move devcoredump generation to a worker")
Acked-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Qiang Yu <Qiang.Yu@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdkfd: Avoid double-unpin of DOORBELL/MMIO BOs on free
Yunxiang Li [Thu, 4 Jun 2026 16:59:11 +0000 (12:59 -0400)]
drm/amdkfd: Avoid double-unpin of DOORBELL/MMIO BOs on free

amdgpu_amdkfd_gpuvm_free_memory_of_gpu() unpinned DOORBELL and MMIO
remap BOs (which are pinned at allocation time) before checking whether
the BO is still mapped to the GPU. When the BO is still mapped, the
function returns -EBUSY and leaves the BO alive, but it has already
been unpinned. The BO is then unpinned again when it is finally freed
during process teardown, triggering a ttm_bo_unpin() underflow warning:

  WARNING: CPU: 18 PID: 15066 at ttm/ttm_bo.c:650 amdttm_bo_unpin+0x6d/0x80 [amdttm]
  Workqueue: kfd_process_wq kfd_process_wq_release [amdgpu]
  RIP: 0010:amdttm_bo_unpin+0x6d/0x80 [amdttm]
  Call Trace:
   amdgpu_bo_unpin+0x1a/0x90 [amdgpu]
   amdgpu_amdkfd_gpuvm_unpin_bo+0x31/0xb0 [amdgpu]
   amdgpu_amdkfd_gpuvm_free_memory_of_gpu+0x3bf/0x460 [amdgpu]
   kfd_process_free_outstanding_kfd_bos+0xd4/0x170 [amdgpu]
   kfd_process_wq_release+0x109/0x1b0 [amdgpu]
   process_one_work+0x1e2/0x3b0
   worker_thread+0x50/0x3a0
   kthread+0xdd/0x100
   ret_from_fork+0x29/0x50

Move the unpin after the mapped_to_gpu_memory check so it only happens
once we are committed to freeing the BO.

Fixes: d25e35bc26c3 ("drm/amdgpu: Pin MMIO/DOORBELL BO's in GTT domain")
Signed-off-by: Yunxiang Li <Yunxiang.Li@amd.com>
Reviewed-by: Felix Kuehling <felix.kuehling@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdkfd: Clean up suspend_all and resume_all mes
Amber Lin [Sat, 30 May 2026 02:25:32 +0000 (22:25 -0400)]
drm/amdkfd: Clean up suspend_all and resume_all mes

Compute user bad/hung queue recovery was handled by KFD using
suspend_all_queues_mes, remove_queue(or reset_queue), and
resume_all_queues_mes. Since now those steps are centralized to
amdgpu_gfx_reset_mes_compute function to sync up with KCQ and KGD user
queues, clean up redundant code and rename the function to match its
functionality.

Signed-off-by: Amber Lin <Amber.Lin@amd.com>
Acked-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu: Disable ras_check_bad_page_status on VFs
Victor Skvortsov [Thu, 4 Jun 2026 13:46:17 +0000 (09:46 -0400)]
drm/amdgpu: Disable ras_check_bad_page_status on VFs

Host driver determines the bad_page_status, not VF.
VFs do not have access to the EEPROM, and eeprom_init
is skipped. However, check_bad_page_status is called
outside of the eeprom_init sequence without any is_vf checks.

Add a return false in __is_ras_eeprom_supported for VFs, and use
that guard in amdgpu_ras_check_bad_page_status to prevent
incorrect access to un-initialized eeprom_control object.

Signed-off-by: Victor Skvortsov <victor.skvortsov@amd.com>
Reviewed-by: Hawking Zhang <Hawking.Zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/pm: Validate OD DPM triples before mutating tables
Asad Kamal [Tue, 2 Jun 2026 18:03:33 +0000 (02:03 +0800)]
drm/amd/pm: Validate OD DPM triples before mutating tables

vega10_odn_edit_dpm_table() and smu7_odn_edit_dpm_table() could mutate
the live ODN table for valid triples, then return 0 after detecting a
truncated buffer or out-of-range index. Validate all (index, clock,
voltage) triples first and return -EINVAL on any failure; only then
apply updates.

v2: Use distinct message for different error case, removed unused
input_level from validation loop (Lijo)

v3: Reject negative level indices, input[] is long but was compared only
against unsigned table bounds, so negative values could pass and truncate
when assigned to uint32_t input_level.

Set DPMTABLE_OD_UPDATE_SCLK/MCLK only after validation passes,
so a failed sysfs write does not leave need_update_dpm_table set for a
later commit.

Signed-off-by: Asad Kamal <asad.kamal@amd.com>
Reviewed-by: Lijo Lazar <lijo.lazar@amd.com>
Reviewed-by: Hawking Zhang <Hawking.Zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/pm: Validate custom profile parameters
Lijo Lazar [Tue, 19 May 2026 11:16:34 +0000 (16:46 +0530)]
drm/amd/pm: Validate custom profile parameters

Add helpers to validate custom profile params against
negative/out-of-range values. Use the helpers to validate user passed
params.

Signed-off-by: Lijo Lazar <lijo.lazar@amd.com>
Assisted-by: Claude Sonnet (Cursor AI)
Reviewed-by: Hawking Zhang <Hawking.Zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu: Gate debugfs MMIO access on kernel lockdown
Asad Kamal [Wed, 3 Jun 2026 09:30:29 +0000 (17:30 +0800)]
drm/amdgpu: Gate debugfs MMIO access on kernel lockdown

amdgpu_regs, amdgpu_regs2, and related debugfs nodes allow
arbitrary MMIO read/write via RREG32/WREG32 without checking
security_locked_down(). On kernel_lockdown=integrity systems
this bypasses the same restrictions as /dev/mem and PCI config
space sysfs.

Check LOCKDOWN_PCI_ACCESS (matching pci-sysfs) at the entry of every
debugfs handler that performs direct register access.

v2: Use consistent check as per previous check to use
LOCKDOWN_DEBUGFS(Lijo)

v3: Do not create any entry from amdgpu_debugfs_regs_init() if
LOCKDOWN_PCI_ACCESS is active and log once. (Lijo)

Signed-off-by: Asad Kamal <asad.kamal@amd.com>
Reviewed-by: Lijo Lazar <lijo.lazar@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu: add ioctl to handle RAS poison error
Yifan Zhang [Wed, 6 May 2026 13:45:05 +0000 (21:45 +0800)]
drm/amdgpu: add ioctl to handle RAS poison error

Add a new DRM_IOCTL_AMDGPU_PROC_OPTIONS ioctl with the
AMDGPU_PROC_OPTIONS_OP_KFD_SIGBUS_DELAY option, allowing userspace (ROCr)
to control per-process SIGBUS delivery.

Userspace for this can be found at:
https://github.com/ROCm/rocm-systems/pull/6190

Reviewed-by: Lijo Lazar <lijo.lazar@amd.com>
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Yifan Zhang <yifan1.zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdkfd: Pass known bad queue info to reset
Amber Lin [Fri, 29 May 2026 21:02:25 +0000 (17:02 -0400)]
drm/amdkfd: Pass known bad queue info to reset

suspend_all, resume_all, and remove bad queue has been integrated to a
centralized function, amdgpu_gfx_reset_mes_compute. Remove remove_queue
and resume_all in KFD and pass the known bad queue information required
for remove_queue to amdgpu_gfx_reset_mes_compute.

Signed-off-by: Amber Lin <Amber.Lin@amd.com>
Acked-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu: Expand MES queue/pipe reset support
Amber Lin [Wed, 6 May 2026 19:02:35 +0000 (15:02 -0400)]
drm/amdgpu: Expand MES queue/pipe reset support

MES in newer versions on gfx11 and gfx12 can support queue/pipe reset via
MES.

v2: update the fw version check (Jesse)

Signed-off-by: Amber Lin <Amber.Lin@amd.com>
Reviewed-by: Jesse Zhang <jesse.zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu: Remove faulty queue before resume
Amber Lin [Fri, 29 May 2026 19:36:52 +0000 (15:36 -0400)]
drm/amdgpu: Remove faulty queue before resume

When driver already knows a bad queue but MES suspend_all is successful
and MES hung queue detection doesn't detect it, remove this queue refore
resume_all.

Signed-off-by: Amber Lin <Amber.Lin@amd.com>
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu/mes12: enable compute MMIO pipe reset
Alex Deucher [Thu, 14 May 2026 19:30:38 +0000 (15:30 -0400)]
drm/amdgpu/mes12: enable compute MMIO pipe reset

Enable MMIO pipe reset for compute pipes.

Reviewed-by: Jesse Zhang <jesse.zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu/mes11: enable compute MMIO pipe reset
Alex Deucher [Thu, 14 May 2026 19:29:29 +0000 (15:29 -0400)]
drm/amdgpu/mes11: enable compute MMIO pipe reset

Enable MMIO pipe reset for compute pipes.

Reviewed-by: Jesse Zhang <jesse.zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu: use a single entry point for mes compute reset
Alex Deucher [Tue, 19 May 2026 22:34:00 +0000 (18:34 -0400)]
drm/amdgpu: use a single entry point for mes compute reset

When we reset MES queues we need to coordinate across
KGD and KFD.  Use a single function to handle the
queue resets across KFD and KGD.

v2: squash in fixes for userqs

Co-developed-by: Jesse Zhang <jesse.zhang@amd.com>
Co-developed-by: Amber Lin <Amber.Lin@amd.com>
Signed-off-by: Amber Lin <Amber.Lin@amd.com>
Signed-off-by: Jesse Zhang <jesse.zhang@amd.com>
Reviewed-by: Jesse Zhang <jesse.zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu/gfx: add a common helper to handle MES compute resets
Alex Deucher [Thu, 7 May 2026 16:03:47 +0000 (12:03 -0400)]
drm/amdgpu/gfx: add a common helper to handle MES compute resets

Add helpers to handle MES compute queue resets when multiple queues
are affected.  Can you be used by both KGD and KFD.

v2: sqaush in updates
v3: squash in userq updates

Co-developed-by: Jesse Zhang <jesse.zhang@amd.com>
Co-developed-by: Amber Lin <Amber.Lin@amd.com>
Signed-off-by: Amber Lin <Amber.Lin@amd.com>
Signed-off-by: Jesse Zhang <jesse.zhang@amd.com>
Reviewed-by: Jesse Zhang <jesse.zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu/userq: add MES userq reset helper
Alex Deucher [Wed, 20 May 2026 20:11:40 +0000 (16:11 -0400)]
drm/amdgpu/userq: add MES userq reset helper

Will be used by the common compute queue reset handler.

Reviewed-by: Jesse Zhang <jesse.zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdkfd: plumb a helper to reset a KFD user queue
Alex Deucher [Mon, 18 May 2026 16:37:19 +0000 (12:37 -0400)]
drm/amdkfd: plumb a helper to reset a KFD user queue

Can be called from KGD.

Reviewed-by: Amber Lin <Amber.Lin@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdkfd: split out mes queue reset sequence into standalone function
Alex Deucher [Mon, 18 May 2026 16:05:15 +0000 (12:05 -0400)]
drm/amdkfd: split out mes queue reset sequence into standalone function

No intended functional change.

Reviewed-by: Amber Lin <Amber.Lin@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu: Use a common KGQ and KCQ reset helper for gfx11/12
Alex Deucher [Tue, 19 May 2026 20:32:59 +0000 (16:32 -0400)]
drm/amdgpu: Use a common KGQ and KCQ reset helper for gfx11/12

They are all the same so use a common implementation.

Reviewed-by: Jesse Zhang <jesse.zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu: store whether to use MMIO or MES for reset
Alex Deucher [Tue, 19 May 2026 21:36:09 +0000 (17:36 -0400)]
drm/amdgpu: store whether to use MMIO or MES for reset

Separate settings for gfx (ME) and compute (MEC).
Use this rather than explicitly specifying it.

Reviewed-by: Jesse Zhang <jesse.zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu/gfx12: unmap the queue via MES on reset for MMIO path
Alex Deucher [Tue, 19 May 2026 20:08:16 +0000 (16:08 -0400)]
drm/amdgpu/gfx12: unmap the queue via MES on reset for MMIO path

To keep MES in sync.

Reviewed-by: Jesse Zhang <jesse.zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu/gfx11: unmap the queue via MES on reset for MMIO path
Alex Deucher [Tue, 19 May 2026 20:06:13 +0000 (16:06 -0400)]
drm/amdgpu/gfx11: unmap the queue via MES on reset for MMIO path

To keep MES in sync.

Reviewed-by: Jesse Zhang <jesse.zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu/gfx12: use the new MQD helper for queue reset
Alex Deucher [Tue, 19 May 2026 20:02:42 +0000 (16:02 -0400)]
drm/amdgpu/gfx12: use the new MQD helper for queue reset

And while we are at it remove the reset parameter as it's
no longer needed.

Reviewed-by: Jesse Zhang <jesse.zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu/gfx11: use the new MQD helper for queue reset
Alex Deucher [Tue, 19 May 2026 19:58:26 +0000 (15:58 -0400)]
drm/amdgpu/gfx11: use the new MQD helper for queue reset

And while we are at it remove the reset parameter as it's
no longer needed.

Reviewed-by: Jesse Zhang <jesse.zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu/gfx: add a helper for MQD restore
Alex Deucher [Tue, 19 May 2026 19:51:53 +0000 (15:51 -0400)]
drm/amdgpu/gfx: add a helper for MQD restore

The handling is common so extract it to a helper.

Reviewed-by: Jesse Zhang <jesse.zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdkfd: rework MES queue reset sequence
Alex Deucher [Thu, 7 May 2026 16:11:29 +0000 (12:11 -0400)]
drm/amdkfd: rework MES queue reset sequence

Call MES with detect only to get the list of hung queues rather
than detecting an resetting.  Then loop over the bad queues
and reset them individually and finally remove them.  Skip
queues not owned by KFD.

v2: always call resume_all after queue reset

Reviewed-by: Amber Lin <Amber.Lin@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu: Allocate enough space for hpd info on gfx11
Amber Lin [Wed, 6 May 2026 19:02:35 +0000 (15:02 -0400)]
drm/amdgpu: Allocate enough space for hpd info on gfx11

MES in newer versions on gfx11 and gfx12 can support queue/pipe
reset via MES.

Signed-off-by: Amber Lin <Amber.Lin@amd.com>
Reviewed-by: Jesse Zhang <jesse.zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amd/amdgpu/include : update mes api header v11/v12
Shaoyun Liu [Mon, 20 Apr 2026 14:45:50 +0000 (10:45 -0400)]
drm/amd/amdgpu/include : update mes api header v11/v12

Update the parameter in SET_HW_RESOURCES API 1. Align with the setting
of enable_lr_compute_wa 2. Add enable_compute_pipe_reset to enable
pipe reset when compute queue reset failes

v2: add driver flags to track when we enable it

Signed-off-by: Shaoyun Liu <shaoyun.liu@amd.com>
Reviewed-by: Jesse Zhang <jesse.zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu/userq: drop detect_and_reset callback
Alex Deucher [Thu, 30 Apr 2026 18:57:59 +0000 (14:57 -0400)]
drm/amdgpu/userq: drop detect_and_reset callback

No longer needed.

Reviewed-by: Jesse Zhang <jesse.zhang@amd.com>
Reviewed-by: Prike Liang <Prike.Liang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
3 months agodrm/amdgpu/userq: switch to per queue reset
Alex Deucher [Wed, 3 Jun 2026 08:38:54 +0000 (16:38 +0800)]
drm/amdgpu/userq: switch to per queue reset

Switch to using the per queue reset rather than
the detect and reset interface.

Reviewed-by: Jesse Zhang <jesse.zhang@amd.com>
Reviewed-by: Prike Liang <Prike.Liang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>