R&D Management 8 min read

Responding to Reviewer Concerns About Hyperparameter Sensitivity: A Practical Rebuttal Guide

This article explains how to effectively respond to reviewer concerns about hyperparameter sensitivity by providing hyperparameter sweeps, justifying default values via validation sets or efficiency trade-offs, and avoiding test-set tuning, with concrete rebuttal templates and examples.

Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Responding to Reviewer Concerns About Hyperparameter Sensitivity: A Practical Rebuttal Guide

In peer review, reviewers often ask whether experimental results are sensitive to hyperparameters — for example, "Why was k set to 40? Are there hyperparameter experiments?" This question probes whether the reported performance depends on a carefully tuned value rather than genuine method robustness.

❌ Not Recommended Response

Many authors reply: "We report the best result the method can achieve, so we chose k=40." This reinforces the reviewer's suspicion that the value was cherry-picked on the test set. Using test-set performance to justify a default hyperparameter risks accusations of tuning on the test set.

Reviewers actually want to know:

Does the method remain effective at other k values?

Is the method sensitive to k?

What is the rationale for k=40?

Was the default chosen on a validation set, from prior work, or by a preset rule?

Is there a performance–efficiency trade-off?

✅ Recommended Response Strategy

Provide a hyperparameter sweep, show stability across a reasonable range, and explain the default choice. A template response:

Thank you for the suggestion. We have reported the sensitivity analysis for k in Table X / Appendix X of the original paper. To avoid oversight, we will emphasize this result in the revision and detail performance changes across k values. Experiments show that with k ∈ {10, 20, 40, 60, 80} the method remains stable and does not rely on a single specific value to achieve effective results. We selected k=40 as the default because it balances validation-set performance and computational efficiency, not because it yields the best test-set score. The revision will further discuss the hyperparameter selection rationale and the performance–efficiency trade-off.

🎯 Three Core Questions to Answer

1. Is the method sensitive to this hyperparameter?

Add a parameter sweep, e.g., k = 10 / 20 / 40 / 60 / 80, and describe the trend:

Across a wide parameter range, the proposed method maintains a consistent advantage.

If performance varies little, state: "The method is not sensitive to this hyperparameter." If there is noticeable variation, explain the cause rather than avoiding it:

When k is too small, XXX information is insufficient; when k is too large, XXX noise is introduced, so the middle range performs best.

Such mechanistic explanations are more convincing.

2. Why was this specific value chosen?

Do not simply say "because it works best." Better justifications include:

Chosen on a validation set.

Follows the default setting of prior work.

Based on a pre-defined rule.

Balances performance and computational cost.

Uses a unified parameter across multiple datasets.

Example: "k=40 achieves a good trade-off between validation performance and computational overhead, so we use it as the default for all experiments."

3. Is there a performance–efficiency trade-off?

Many hyperparameters affect not only accuracy but also:

Inference time

GPU memory usage

Retrieval count

Token count

API call count

Training or inference cost

Therefore, explain the trade-off explicitly:

Increasing k from 40 to 80 yields diminishing performance gains while significantly increasing computational cost, so we finalize k=40.

This reasoning is more principled than reporting only the peak.

⚠️ Critical Warning

If you actually selected k=40 based on test-set results, do not disguise it as a principled choice in the rebuttal. The correct remedy is to re-determine the hyperparameter on a validation set, fix it, and re-report test-set results. Hyperparameters must be chosen on the validation set; the test set is reserved for final evaluation only.

📌 One-Sentence Summary

Hyperparameter sweeps should demonstrate method stability and justify the default value, not merely report the best result. Focus on the tested range, observed trend, and why the final choice is reasonable.

Project repository: https://github.com/MLNLP-World/Paper-Rebuttal-Tips

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

machine learningacademic writinghyperparameter tuningpeer reviewvalidation settest setpaper rebuttalhyperparameter sensitivity
Machine Learning Algorithms & Natural Language Processing
Written by

Machine Learning Algorithms & Natural Language Processing

Focused on frontier AI technologies, empowering AI researchers' progress.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.