주요 콘텐츠

Match Data to Designs

R2026b

After you have collected data against a Design of Experiments test plan, you can use the Design Match view to see how well your data aligns with the planned test points. In this view, you can match your collected data to the experimental design, select the data you want to use for modeling, and assess how closely it follows the intended design structure. This view supports an iterative workflow: create a design, collect data, match the data to the design points, update the design based on the data you have collected, and then gather additional data as needed. By repeating these steps, you can optimize your data collection process to build robust global models while minimizing the amount of data required.

All data you select using the Design Match view is added to a new design called Actual Design. You can use the matching process to produce an Actual Design that accurately reflects your current data. You can then use this new design to decide the best points to use if you want to augment your current design in order to collect more data.

How to Use the Design Match View

When you open the Data Editor from an existing test plan, the Match data to design button option is available. To zoom in on clusters of interest in the Design Match plot, press Shift+click and drag. Double-click the plot to return to full size.

To match data to designs using the Design Match plot:

  1. Using the context menu, open the Tolerance Editor using the context menu. Try different values for different variables. You will likely need to experiment to find the right tolerances. These values determine the size of clusters centered on each design point. Data points that lie within tolerance of any design point in a cluster are matched to that cluster. For cluster definitions, see The Tolerance Editor.

  2. For matching data to designs, clear the check box in the Design Match plot for green clusters (with equal data and design points). Green clusters are matched, so removing them allows you to focus unmatched points and clusters with uneven numbers of data and design points. If you want your new Actual Design to accurately reflect your current data, you should try to get as many data points matched to design points as possible, which means you want as few red clusters as possible. See Red Clusters.

  3. See the values of variables at different points by clicking and holding. Selected points have a pink border. Once you select points, you can change the plot variables using the X-axis factor and Y-axis factor lists to track those points through the different dimensions.

    Tracking can help you to decide which tolerances to change in order to match points. Remember that points that do not form a cluster can appear to be perfectly matched when viewed in one pair of dimensions. You must view points in other dimensions to find out where they are separated beyond the tolerance value. Use this tracking process to decide whether you want particular pairs of points to be matched, and then change the tolerances until they form part of a cluster.

    The points you select in the design match view are selected across the Data Editor, so if you have other data plots or a table view open, you can investigate the same points in different views.

  4. Once you find useful values for the tolerances, you can select points within clusters that have uneven numbers of data and design points. These clusters are blue (more data than design) or red (more design than data). Select any cluster by clicking it. The details of every data and design point contained in the selected cluster appear in the Cluster Information list. Choose the points you want to keep or discard by selecting or clearing the check box next to each point. Your selections can cause clusters to change color as you adjust the numbers of data and design points within them.

  5. You can also select unmatched points by right-clicking and selecting Select Unmatched Data. All unmatched points then appear in the list view. You can decide whether to include or exclude them in the same way as points within clusters, by using the check boxes in the list. If you decide to exclude data points (within clusters or not), they appear on the plot as black crosses if the Excluded Data check box is selected for display.

    Selecting multiple points before selecting or clearing a check box is faster than selecting points individually. To select multiple points, press Shift+click, and then hold the Shift key when selecting one of the check boxes.

    You can right-click and select Show Labels to see design and data point numbers on the plot (also in the View menu).

Continue altering tolerances and selecting points until you have selected all the data you want for modeling. All selected data is also added to your new Actual Design, in red clusters.

Red Clusters

Red clusters contain more design points than data points. These data points are not added to your design because the algorithm cannot choose the design points to replace, so you must make manual selections to deal with red clusters if you want to use these data points in your design. If you are just selecting data for modeling without regard for the Actual Design, then you can ignore red clusters. The data points in red clusters are selected for modeling. For information about the effects of your selections, see What Will Happen to My Data and Design?

The Tolerance Editor

Open the Tolerance Editor by selecting Tolerances in the context menu.

In the Tolerance Editor, you can edit the tolerance for selecting data points. You can choose values for each variable to determine the size of tolerance in each dimension.

  • Data points within the tolerance of a design point are included in that cluster.

  • Data points that fall inside the tolerance of more than one design point form a single cluster containing all those design and data points.

  • Excluded data (shown as black crosses) that lies within tolerance appears in the list when that cluster is selected. You can then choose whether to use it or continue to exclude it.

  • Data in Design (pink crosses) is the only type of data that is not included in clusters.

    Note

    For grouped data, tolerances are set for global variables. Data used for matching uses operating point means of global variables, not individual records, unlike other Data Editor views. Select points to inspect values of global variables.

Using the Tolerance Editor is the same process as setting tolerances within the Data Wizard. In the Data Wizard, you can also choose in advance what to do with unmatched data and clusters with uneven numbers of data and design points. These choices affect how the cluster algorithm is first run, though you can change selections later in the Data Editor. See Step 4: Set Tolerances.

Note

If you modify the data in any way while the Design Match view is open, such as applying a filter, the cluster algorithm will be rerun. You might lose your design point selections.

What Will Happen to My Data and Design?

The changes you make in the Design Match plot are only applied to the data set when you exit the Data Editor. When you close the Data Editor, a new design called Actual Design is created. All the changes are determined by your check box selections for data and design points.

Note

All data points with a selected check box are selected for modeling. All data points with a cleared check box are excluded from the data set.

All data points with a selected check box are put into the new Actual Design except those in red clusters.

When you close the Data Editor, these changes are applied:

  • Green clusters—Equal number of data and design points. The design points are replaced by the equal number of data points. These points become fixed design points (red in the Design Editor table) and appear as Data in Design (pink crosses) when you reopen the Data Editor.

    This means that these points are not included in clusters when matching again. These fixed points are also not changed in the Design Editor when you add points, although you can unlock fixed points if you want. This can be very useful if you want to optimally augment a design, taking into account the data you have already obtained.

  • Blue clusters—More data than design points. The design points are replaced by all the data points.

    Note

    Design points with selected check boxes in green or blue clusters are the points that will be replaced by your selected data points. You may have cleared the check boxes of other design points in these clusters, and these points will be left unchanged.

  • Red clusters— More design than data points. Red clusters indicate that you should make a decision if you want your new Actual Design to reflect your most current data. The algorithm cannot choose the design points to replace with the data points, so no action is taken. Red clusters do not make any changes to the design when you close the data editor. The existing design points remain in the design. The data points are included or excluded from the data set depending on your selections in the Cluster Information list, but they are not added to the design.

  • Unmatched Design Points—These remain in the design.

  • Unmatched Data Points—If you have selected the check boxes for unmatched data, they become new fixed design points, which are red in the Design Editor. When you reopen the Data Editor, these points are Data in Design, which appear as pink crosses. Note that in the Data Wizard you could choose Use to select all these initially, or you could choose Do not use, which clears all their check boxes. See Step 4: Set Tolerances.

  • Data in Design—These remain in the design.

  • Excluded Data—These data points are removed from the data set and are not displayed in any other views. If you want to return them to the data set you can only do so by selecting them in the Design Match view.