Quantum architecture search (QAS) has shown significant potential for automating quantum circuit design in variational quantum algorithms (VQAs). However, the application of current QAS methods to large-scale circuits remains challenging due to the high computational complexity of evaluating numerous circuit architectures. While performance predictors are a promising solution, they are conventionally trained with objectives such as mean squared error, which create a fundamental objective mismatch by failing to prioritize the discovery of the most elite, high-performing architectures. In this work, we resolve this objective mismatch by introducing a new training paradigm that directly optimizes a top-tier-focused ranking loss. To capture the causal information flow of quantum gates, we design a circuit-native Transformer equipped with reachability-based unitary-flow attention and depth-aware positional encoding. Furthermore, we implement a pretraining strategy based on circuit learnability, a metric that unifies expressibility and trainability, to extract functionally aware representations from unlabeled data, thereby significantly improving sample efficiency. Extensive results on multiple VQA benchmarks demonstrate that our approach consistently outperforms existing QAS frameworks in both ranking quality and the performance of the discovered circuits. Ablation studies confirm that the synergy of the differentiable top-tier-focused loss, the circuit-native Transformer, and learnability-based pretraining is essential for superior performance. Beyond energy-based benchmarks, we further validate the searched circuits under a device-calibrated noise model based on the IBM Kingston quantum computer and on a noisy 2D transverse-field Ising model benchmark. The results show that our method can identify compact circuits with fewer two-qubit gates while preserving physically meaningful properties, including fidelity, connected spin correlations, and magnetization.